MiniMax-H3 video decode only enables fp16 autocast on CUDA, so the block silently runs in fp32 on other accelerators
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 1/5
- Thời gian dự kiến
- Dưới một giờ
- Mức phù hợp với người mới
- 86/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Lĩnh vực
- machine-learning
Hướng nghiên cứu
Bắt đầu tại src/diffusers/modular_pipelines/minimax_h3/decoders.py, ở MiniMaxH3VideoDecodeStep.call, và kiểm tra autocast gate quanh phần giải mã VAE. Xác minh rằng các bộ tăng tốc không phải CPU sử dụng autocast với float16, CPU vẫn bị vô hiệu hóa và hành vi của CUDA không thay đổi; việc tái hiện dtype của issue cung cấp một cách để kiểm tra đường dẫn dành riêng cho thiết bị.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the bug
MiniMaxH3VideoDecodeStep.__call__ (src/diffusers/modular_pipelines/minimax_h3/decoders.py, line ~187) wraps the VAE decode in an fp16 autocast whose gate is CUDA-only:
with torch.autocast(device_type=device.type, dtype=torch.float16, enabled=device.type == "cuda"):
video = components.vae.decode(latents, return_dict=False)[0]
The block's own description states the intent - "the decode itself runs under float16 autocast even though the VAE weights are float32" - but with enabled=device.type == "cuda" that is only true on CUDA. On every other accelerator (Ascend NPU, Intel XPU, Apple MPS) autocast is silently disabled, so the same workflow decodes with the VAE's fp32 weights: slower, and numerically on a different path than the documented CUDA behaviour.
Suggested change - gate on "not CPU" instead of "CUDA only":
enabled=device.type != "cpu",
CPU keeps the current behaviour (autocast disabled).
Reproduction
On Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu + torch_npu), the enabled flag is the only thing keeping the NPU on the fp32 path:
import torch
x = torch.randn(8, 8, device="npu", dtype=torch.float32)
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=True):
print((x @ x).dtype) # torch.float16
with torch.autocast(device_type="npu", dtype=torch.float16, enabled=False):
print((x @ x).dtype) # torch.float32
Not measured
End-to-end NPU decode speed-up of the change is not measured here - no A/B benchmark of the decode block was run, only the dtype-path check above. CUDA behaviour is unchanged by the linked PR.
Related but distinct
#14746 reports a trade-off in the opposite direction on CUDA (fp32 VAE weights + fp16 decode costs VRAM and forces a downcast). This issue is only about the device gate being CUDA-only; it does not argue about whether the fp16 autocast should exist at all.
System Info
- diffusers: main
- hardware: Ascend 910B2 NPU (torch 2.15.0.dev20260917+cpu +
torch_npu); applies to any non-CUDA accelerator.
- Ngôn ngữ chính
- Python
- Star
- 34.6k
- Fork
- 7.4k
- Merge trung bình
- 4 ngày 11 giờ
- Pull request đã merge (30 ngày)
- 48
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của huggingface/diffusers
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
huggingface/diffusers#14888 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
huggingface/diffusers#14881 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
huggingface/diffusers#14864 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
huggingface/diffusers#14837 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug needs-env-info pipelines
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
huggingface/diffusers#14794 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của huggingface/diffusers
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
spec-kitty/spec-kitty#5319 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
backend::vllm diffusion multimodal
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
openai/openai-agents-python#5229 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày