Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

MiniMax-H3 video decode only enables fp16 autocast on CUDA, so the block silently runs in fp32 on other accelerators

已关闭 适合新手
#14,882 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
1/5
预计耗时
1 小时以内
新手友好度
86/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
python, pytorch

调研方向

从 src/diffusers/modular_pipelines/minimax_h3/decoders.py 中的 MiniMaxH3VideoDecodeStep.call 开始,检查 VAE decode 周围的 autocast gate。确认非 CPU 加速器使用 float16 autocast,CPU 保持禁用,并且 CUDA 行为不变;issue 的 dtype 复现提供了一种检查设备特定路径的方法。

由索引模型根据 Issue 内容生成。

描述

Describe the bug

MiniMaxH3VideoDecodeStep.__call__ (src/diffusers/modular_pipelines/minimax_h3/decoders.py, line ~187) wraps the VAE decode in an fp16 autocast whose gate is CUDA-only:

with torch.autocast(device_type=device.type, dtype=torch.float16, enabled=device.type == "cuda"):
    video = components.vae.decode(latents, return_dict=False)[0]

The block's own description states the intent - "the decode itself runs under float16 autocast even though the VAE weights are float32" - but with enabled=device.type == "cuda" that is only true on CUDA. On every other accelerator (Ascend NPU, Intel XPU, Apple MPS) autocast is silently disabled, so the same workflow decodes with the VAE's fp32 weights: slower, and numerically on a different path than the documented CUDA behaviour.

Suggested change - gate on "not CPU" instead of "CUDA only":

enabled=device.type != "cpu",

CPU keeps the current behaviour (autocast disabled).

Reproduction

On Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu), the enabled flag is the only thing keeping the NPU on the fp32 path:

import torch

x = torch.randn(8, 8, device="npu", dtype=torch.float32)
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=True):
    print((x @ x).dtype)   # torch.float16
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=False):
    print((x @ x).dtype)   # torch.float32
Not measured

End-to-end NPU decode speed-up of the change is not measured here - no A/B benchmark of the decode block was run, only the dtype-path check above. CUDA behaviour is unchanged by the linked PR.

Related but distinct

#14746 reports a trade-off in the opposite direction on CUDA (fp32 VAE weights + fp16 decode costs VRAM and forces a downcast). This issue is only about the device gate being CUDA-only; it does not argue about whether the fp16 autocast should exist at all.

System Info
  • diffusers: main
  • hardware: Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu); applies to any non-CUDA accelerator.
主要语言
Python
星标
34.6k
派生
7.4k
平均合并
4 天 8 小时
30 天内合并 PR
50

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

huggingface/diffusers 的其他 Issue

查看 huggingface/diffusers 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。