Distributed gather crashes on PyTorch < 2.6
维护者通常 1 天内回复
@akshatvishu 已经在做这个了。
开始于 2026年2月3日。
评估
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 新手友好度
- 68/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 停滞
调研方向
从 src/diffusers/models/_modeling_parallel.py 中的 gather_size_by_comm 开始,并使用 PyTorch 2.4 运行 issue 中的最小分布式复现。跟踪通信后端如何确定 gather_device。完成标准是:在受支持的 PyTorch 2.1–2.5 上,分布式 gather 不再引发 AttributeError,同时保留在更新版本上的行为。
由索引模型根据 Issue 内容生成。
描述
Describe the bug
AttributeError: module 'torch' has no attribute 'accelerator' when running distributed gather on PyTorch versions < 2.6.
This error happens because gather_size_by_comm in src/diffusers/models/_modeling_parallel.py uses torch.accelerator.current_accelerator(), which only exists in PyTorch 2.6+. Diffusers officially supports PyTorch 2.1+, so this causes a crash on versions 2.1–2.5 with AttributeError: module 'torch' has no attribute 'accelerator'.
Reproduction
Since this is a utility function; it can be triggered directly with a minimal distributed setup:
import torch.distributed as dist
from diffusers.models._modeling_parallel import gather_size_by_comm
dist.init_process_group(
backend="gloo",
init_method="file:///tmp/pg",
rank=0,
world_size=1,
)
gather_size_by_comm(1, dist.group.WORLD)
Logs
[rank0]: Traceback (most recent call last):
[rank0]: File "/home/aja/diffusers/test.py", line 11, in <module>
[rank0]: gather_size_by_comm(1, dist.group.WORLD)
[rank0]: File "/home/aja/diffusers/src/diffusers/models/_modeling_parallel.py", line 293, in gather_size_by_comm
[rank0]: gather_device = "cpu" if "cpu" in comm_backends else torch.accelerator.current_accelerator()
[rank0]: ^^^^^^^^^^^^^^^^^
[rank0]: File "/home/aja/diffusers/.venv/lib/python3.11/site-packages/torch/__init__.py", line 2216, in __getattr__
[rank0]: raise AttributeError(f"module '{__name__}' has no attribute '{name}'")
[rank0]: AttributeError: module 'torch' has no attribute 'accelerator'
System Info
diffuser : 0.37.0.dev0
torch : 2.4.0
python : 3.11
system: Linux
Who can help?
@sayakpaul @DN6
- 主要语言
- Python
- 星标
- 34.6k
- 派生
- 7.4k
- 平均合并
- 4 天 16 小时
- 30 天内合并 PR
- 44
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
huggingface/diffusers 的其他 Issue
-
bug models
难度 2/5 1-3 小时 新手友好度 70/100
huggingface/diffusers#14952 ·
维护者通常 1 天内回复
-
UniPCMultistepScheduler fails in torch.linalg.solve under a float64 default dtype: the unit entry of rks takes the default dtype可能已有人在做 @DawnofGenX 于 5 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
huggingface/diffusers#14888 ·
维护者通常 1 天内回复
-
WanAnimatePipeline.get_i2v_mask() defaults device to "cuda", which raises on non-CUDA accelerators (NPU/XPU/MPS)可能已有人在做 @li-lizhe 于 8 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
huggingface/diffusers#14881 ·
维护者通常 1 天内回复
-
IndexError when preprocessing an empty image or video list可能已有人在做 @MohammadHijjawi97 于 13 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 74/100
huggingface/diffusers#14864 ·
维护者通常 1 天内回复
-
Windows: check_ai.py fails decoding UTF-8 guides with the default locale可能已有人在做 @tanvir-ux 于 13 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 84/100
huggingface/diffusers#14837 ·
维护者通常 1 天内回复
查看 huggingface/diffusers 的全部 Issue
相似的 Issue
-
Device Details tables: FS/SF columns contradict each other (nfet_01v8 Vt row, pfet_01v8 Idsat row)未关闭
难度 2/5 1-3 小时 新手友好度 75/100
google/skywater-pdk#450 ·
-
Drained trajectory arrays are overwritten when the sequence buffer is reused可能已有人在做 @sylvesterkaczmarek 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 78/100
google-deepmind/bsuite#56 ·
-
难度 2/5 1-3 小时 新手友好度 82/100
LearningCircuit/local-deep-research#7206 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 68/100
chingu-voyages/V62-tier3-team-33#285 ·
维护者通常 1 天内回复
-
Proxy drops log notifications from backends that don't send FastMCP's msg/extra dict可能已有人在做 @asasemahmed 今天认领。 未关闭bug server
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复