Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

MiniMax-H3 video decode only enables fp16 autocast on CUDA, so the block silently runs in fp32 on other accelerators

オープン 初心者向け
#14,882 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
1/5
見積もり時間
1時間未満
初心者へのやさしさ
86/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
python, pytorch

調査の方向性

src/diffusers/modular_pipelines/minimax_h3/decoders.py の MiniMaxH3VideoDecodeStep.call から始め、VAE decode 周辺の autocast ゲートを確認します。CPU 以外のアクセラレータでは float16 autocast が使用され、CPU では無効のままであり、CUDA の動作は変更されないことを検証してください。issue の dtype 再現により、デバイス固有のパスを確認できます。

索引モデルが issue の本文から書いたものです。

説明

Describe the bug

MiniMaxH3VideoDecodeStep.__call__ (src/diffusers/modular_pipelines/minimax_h3/decoders.py, line ~187) wraps the VAE decode in an fp16 autocast whose gate is CUDA-only:

with torch.autocast(device_type=device.type, dtype=torch.float16, enabled=device.type == "cuda"):
    video = components.vae.decode(latents, return_dict=False)[0]

The block's own description states the intent - "the decode itself runs under float16 autocast even though the VAE weights are float32" - but with enabled=device.type == "cuda" that is only true on CUDA. On every other accelerator (Ascend NPU, Intel XPU, Apple MPS) autocast is silently disabled, so the same workflow decodes with the VAE's fp32 weights: slower, and numerically on a different path than the documented CUDA behaviour.

Suggested change - gate on "not CPU" instead of "CUDA only":

enabled=device.type != "cpu",

CPU keeps the current behaviour (autocast disabled).

Reproduction

On Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu), the enabled flag is the only thing keeping the NPU on the fp32 path:

import torch

x = torch.randn(8, 8, device="npu", dtype=torch.float32)
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=True):
    print((x @ x).dtype)   # torch.float16
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=False):
    print((x @ x).dtype)   # torch.float32
Not measured

End-to-end NPU decode speed-up of the change is not measured here - no A/B benchmark of the decode block was run, only the dtype-path check above. CUDA behaviour is unchanged by the linked PR.

Related but distinct

#14746 reports a trade-off in the opposite direction on CUDA (fp32 VAE weights + fp16 decode costs VRAM and forces a downcast). This issue is only about the device gate being CUDA-only; it does not argue about whether the fp16 autocast should exist at all.

System Info
  • diffusers: main
  • hardware: Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu); applies to any non-CUDA accelerator.
主要言語
Python
スター
34.6k
フォーク
7.4k
平均マージ
4日 11時間
マージ済み PR(30日)
48

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

huggingface/diffusers のほかの issue

huggingface/diffusers の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。