MiniMax-H3 video decode only enables fp16 autocast on CUDA, so the block silently runs in fp32 on other accelerators
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 1/5
- 预计耗时
- 1 小时以内
- 新手友好度
- 86/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 活跃
调研方向
从 src/diffusers/modular_pipelines/minimax_h3/decoders.py 中的 MiniMaxH3VideoDecodeStep.call 开始,检查 VAE decode 周围的 autocast gate。确认非 CPU 加速器使用 float16 autocast,CPU 保持禁用,并且 CUDA 行为不变;issue 的 dtype 复现提供了一种检查设备特定路径的方法。
由索引模型根据 Issue 内容生成。
描述
Describe the bug
MiniMaxH3VideoDecodeStep.__call__ (src/diffusers/modular_pipelines/minimax_h3/decoders.py, line ~187) wraps the VAE decode in an fp16 autocast whose gate is CUDA-only:
with torch.autocast(device_type=device.type, dtype=torch.float16, enabled=device.type == "cuda"):
video = components.vae.decode(latents, return_dict=False)[0]
The block's own description states the intent - "the decode itself runs under float16 autocast even though the VAE weights are float32" - but with enabled=device.type == "cuda" that is only true on CUDA. On every other accelerator (Ascend NPU, Intel XPU, Apple MPS) autocast is silently disabled, so the same workflow decodes with the VAE's fp32 weights: slower, and numerically on a different path than the documented CUDA behaviour.
Suggested change - gate on "not CPU" instead of "CUDA only":
enabled=device.type != "cpu",
CPU keeps the current behaviour (autocast disabled).
Reproduction
On Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu), the enabled flag is the only thing keeping the NPU on the fp32 path:
import torch
x = torch.randn(8, 8, device="npu", dtype=torch.float32)
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=True):
print((x @ x).dtype) # torch.float16
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=False):
print((x @ x).dtype) # torch.float32
Not measured
End-to-end NPU decode speed-up of the change is not measured here - no A/B benchmark of the decode block was run, only the dtype-path check above. CUDA behaviour is unchanged by the linked PR.
Related but distinct
#14746 reports a trade-off in the opposite direction on CUDA (fp32 VAE weights + fp16 decode costs VRAM and forces a downcast). This issue is only about the device gate being CUDA-only; it does not argue about whether the fp16 autocast should exist at all.
System Info
- diffusers: main
- hardware: Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu +
torch_npu); applies to any non-CUDA accelerator.
- 主要语言
- Python
- 星标
- 34.6k
- 派生
- 7.4k
- 平均合并
- 4 天 8 小时
- 30 天内合并 PR
- 50
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
huggingface/diffusers 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 88/100
huggingface/diffusers#14888 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
huggingface/diffusers#14881 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 74/100
huggingface/diffusers#14864 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
huggingface/diffusers#14837 ·
维护者通常 1 天内回复
-
bug needs-env-info pipelines
难度 2/5 1-3 小时 新手友好度 72/100
huggingface/diffusers#14794 ·
维护者通常 1 天内回复
查看 huggingface/diffusers 的全部 Issue
相似的 Issue
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)未关闭
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattended未关闭workflow
难度 2/5 1-3 小时 新手友好度 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
维护者通常 1 天内回复
-
metadata submission
难度 2/5 1-3 小时 新手友好度 82/100
-
bug
难度 2/5 1-3 小时 新手友好度 65/100
canonical/content-cache-operator#163 · 1 条评论 ·
维护者通常 1 天内回复
-
[submission]未关闭submission
难度 1/5 1 小时以内 新手友好度 65/100
leanprover/lean-eval-submissions#1852 ·
维护者通常 1 天内回复