Matrix-Game-3: all multi-GPU runs fail with TypeError - attention() called with fa_version= instead of version=
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 1/5
- Tempo stimato
- Meno di un'ora
- Idoneità per principianti
- 88/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Tranquilla
- Ambito
- distributed-systems, machine-learning
Direzione di ricerca
Inizia da Matrix-Game-3/wan/distributed/ulysses.py:63-70 e confronta la keyword di attention() con la firma in wan/modules/attention.py:176. Esegui la riproduzione multi-GPU fornita oppure Matrix-Game-3/test.sh; il lavoro è completato quando il percorso documentato con --ulysses_size maggiore di 1 non solleva più un TypeError unexpected-keyword.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
Every run with --ulysses_size > 1 crashes immediately in the first DiT block with a TypeError. This includes the configuration in the repo's own Matrix-Game-3/test.sh (which uses --ulysses_size 7/8), so the documented multi-GPU path does not run as shipped.
Tested at commit 71c3cd7f741311f8100f6cf9cde942b6c1378d11.
Error
File "Matrix-Game-3/wan/distributed/ulysses.py", line 63, in distributed_attention
x = attention(
TypeError: attention() got an unexpected keyword argument 'fa_version'
Cause
wan/distributed/ulysses.py:63-70 calls attention() with fa_version=:
from ..modules.attention import attention
x = attention(
q, k, v,
k_lens=seq_lens,
window_size=window_size,
fa_version=fa_version, # <-- wrong keyword
)
but the parameter in wan/modules/attention.py:176 is named version:
def attention(
q, k, v,
...
version=None,
):
Only this call site is affected. The other fa_version=fa_version call sites (sequence_parallel.py:307,364 and model.py:591,596,613,615,678,1015,1027) target functions that do declare fa_version, so they are correct.
Fix
One word, in wan/distributed/ulysses.py:69:
- fa_version=fa_version,
+ version=fa_version,
Reproduce
torchrun --nproc_per_node=2 generate.py \
--size 704*1280 --ckpt_dir Matrix-Game-3.0 \
--image demo_images/001/image.png --prompt "..." \
--ulysses_size 2 --dit_fsdp --t5_fsdp \
--num_iterations 3 --num_inference_steps 3 --use_int8
Fails identically regardless of --fa_version, GPU model, or whether FlashAttention is installed — the TypeError is raised before any attention backend is selected.
Environment
2x RTX 5090 (sm_120), torch 2.11.0+cu128, Python 3.12, WSL2 Ubuntu 22.04. The bug is hardware-independent.
Related minor issues
Happy to open these separately if you'd prefer — noting them here since they came up while getting the model running on 2 GPUs:
-
WAN_VAE_SEGMENT_SIZEis undocumented.inference_pipeline.py:643reads it (default4), and it decides whether inference fits in 32 GB. At4the VAE decode exhausts VRAM at 704x1280 even with--use_int8,--t5_cpuand--lightvae_pruning_rate 0.5; at1the same run completes. A README note or a--vae_segment_sizeflag would help. -
stream_decodeswallows exceptions.wan/modules/vae2_2.py:1349catchesException, logs one line, and returnsNone. ThatNonethen surfaces ~200 lines away asAttributeError: 'NoneType' object has no attribute 'cpu'(inference_pipeline.py:652), hiding the real CUDA/OOM error. Addingtraceback.format_exc()there would make debugging much easier.
Thanks for open-sourcing the weights and code — the one-word fix above got 2-GPU inference working end to end for us.
- Lingua principale
- Python
- Stelle
- 2.3k
- Fork
- 256
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di SkyworkAI/Matrix-Game
-
Replace deprecated v2.ToTensor() in Matrix-Game-2/inference.py with v2.ToImage() + v2.ToDtype(scale=True)Forse già presa @xyf5432 l’ha presa 48 giorni fa. Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
SkyworkAI/Matrix-Game#82 ·
-
run logic error when finishForse già presa @Arison591 l’ha presa 85 giorni fa. Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
SkyworkAI/Matrix-Game#77 ·
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 76/100
SkyworkAI/Matrix-Game#86 ·
-
Third person supportAperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
SkyworkAI/Matrix-Game#76 · 1 commento ·
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 25/100
SkyworkAI/Matrix-Game#75 · 1 reazione ·
Tutte le issue di SkyworkAI/Matrix-Game
Issue simili
-
enhancement good first issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
-
python-version
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
I maintainer di solito rispondono entro 1 giorno
-
bug javascript P2-medium python release:v3.1
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
adrirubio/claude-deck#546 ·
I maintainer di solito rispondono entro 1 giorno
-
area: desktop area: website priority: P2 type: feature
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
appandflow/stim#3411 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno