MiniMax-H3 video decode only enables fp16 autocast on CUDA, so the block silently runs in fp32 on other accelerators
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 1/5
- 見積もり時間
- 1時間未満
- 初心者へのやさしさ
- 86/100
- issue の種類
- バグ
- 明瞭さ
- 明確に書かれている
- 活発さ
- 活発
調査の方向性
src/diffusers/modular_pipelines/minimax_h3/decoders.py の MiniMaxH3VideoDecodeStep.call から始め、VAE decode 周辺の autocast ゲートを確認します。CPU 以外のアクセラレータでは float16 autocast が使用され、CPU では無効のままであり、CUDA の動作は変更されないことを検証してください。issue の dtype 再現により、デバイス固有のパスを確認できます。
索引モデルが issue の本文から書いたものです。
説明
Describe the bug
MiniMaxH3VideoDecodeStep.__call__ (src/diffusers/modular_pipelines/minimax_h3/decoders.py, line ~187) wraps the VAE decode in an fp16 autocast whose gate is CUDA-only:
with torch.autocast(device_type=device.type, dtype=torch.float16, enabled=device.type == "cuda"):
video = components.vae.decode(latents, return_dict=False)[0]
The block's own description states the intent - "the decode itself runs under float16 autocast even though the VAE weights are float32" - but with enabled=device.type == "cuda" that is only true on CUDA. On every other accelerator (Ascend NPU, Intel XPU, Apple MPS) autocast is silently disabled, so the same workflow decodes with the VAE's fp32 weights: slower, and numerically on a different path than the documented CUDA behaviour.
Suggested change - gate on "not CPU" instead of "CUDA only":
enabled=device.type != "cpu",
CPU keeps the current behaviour (autocast disabled).
Reproduction
On Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu), the enabled flag is the only thing keeping the NPU on the fp32 path:
import torch
x = torch.randn(8, 8, device="npu", dtype=torch.float32)
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=True):
print((x @ x).dtype) # torch.float16
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=False):
print((x @ x).dtype) # torch.float32
Not measured
End-to-end NPU decode speed-up of the change is not measured here - no A/B benchmark of the decode block was run, only the dtype-path check above. CUDA behaviour is unchanged by the linked PR.
Related but distinct
#14746 reports a trade-off in the opposite direction on CUDA (fp32 VAE weights + fp16 decode costs VRAM and forces a downcast). This issue is only about the device gate being CUDA-only; it does not argue about whether the fp16 autocast should exist at all.
System Info
- diffusers: main
- hardware: Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu +
torch_npu); applies to any non-CUDA accelerator.
- 主要言語
- Python
- スター
- 34.6k
- フォーク
- 7.4k
- 平均マージ
- 4日 11時間
- マージ済み PR(30日)
- 48
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
huggingface/diffusers のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
huggingface/diffusers#14888 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
huggingface/diffusers#14881 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
huggingface/diffusers#14864 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
huggingface/diffusers#14837 ·
メンテナーはふだん 1 日以内に返信
-
bug needs-env-info pipelines
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
huggingface/diffusers#14794 ·
メンテナーはふだん 1 日以内に返信
huggingface/diffusers の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜2日 初心者へのやさしさ 70/100
-
FingerprintSplitter raises ZeroDivisionError when int(frac_train * len(dataset)) floors to zeroオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
メンテナーはふだん 7 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
lmstudio-ai/mlx-engine#376 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
pyiron/bagofholding#166 ·