Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

MiniMax-H3 video decode only enables fp16 autocast on CUDA, so the block silently runs in fp32 on other accelerators

Chiusa Adatta ai principianti
#14,882 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
1/5
Tempo stimato
Meno di un'ora
Idoneità per principianti
86/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python, pytorch

Direzione di ricerca

Inizia in src/diffusers/modular_pipelines/minimax_h3/decoders.py, in MiniMaxH3VideoDecodeStep.call, e analizza il gate di autocast attorno alla decodifica VAE. Verifica che gli acceleratori diversi da CPU usino l’autocast con float16, che CPU rimanga disabilitato e che il comportamento di CUDA non cambi; la riproduzione del dtype dell’issue offre un modo per verificare il percorso specifico del dispositivo.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Describe the bug

MiniMaxH3VideoDecodeStep.__call__ (src/diffusers/modular_pipelines/minimax_h3/decoders.py, line ~187) wraps the VAE decode in an fp16 autocast whose gate is CUDA-only:

with torch.autocast(device_type=device.type, dtype=torch.float16, enabled=device.type == "cuda"):
    video = components.vae.decode(latents, return_dict=False)[0]

The block's own description states the intent - "the decode itself runs under float16 autocast even though the VAE weights are float32" - but with enabled=device.type == "cuda" that is only true on CUDA. On every other accelerator (Ascend NPU, Intel XPU, Apple MPS) autocast is silently disabled, so the same workflow decodes with the VAE's fp32 weights: slower, and numerically on a different path than the documented CUDA behaviour.

Suggested change - gate on "not CPU" instead of "CUDA only":

enabled=device.type != "cpu",

CPU keeps the current behaviour (autocast disabled).

Reproduction

On Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu), the enabled flag is the only thing keeping the NPU on the fp32 path:

import torch

x = torch.randn(8, 8, device="npu", dtype=torch.float32)
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=True):
    print((x @ x).dtype)   # torch.float16
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=False):
    print((x @ x).dtype)   # torch.float32
Not measured

End-to-end NPU decode speed-up of the change is not measured here - no A/B benchmark of the decode block was run, only the dtype-path check above. CUDA behaviour is unchanged by the linked PR.

Related but distinct

#14746 reports a trade-off in the opposite direction on CUDA (fp32 VAE weights + fp16 decode costs VRAM and forces a downcast). This issue is only about the device gate being CUDA-only; it does not argue about whether the fp16 autocast should exist at all.

System Info
  • diffusers: main
  • hardware: Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu); applies to any non-CUDA accelerator.
Lingua principale
Python
Stelle
34.6k
Fork
7.4k
Merge medio
4g 8h
PR unite (30g)
50

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di huggingface/diffusers

Tutte le issue di huggingface/diffusers

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.