Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Matrix-Game-3: all multi-GPU runs fail with TypeError - attention() called with fa_version= instead of version=

Aperta Adatta ai principianti
#80 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
1/5
Tempo stimato
Meno di un'ora
Idoneità per principianti
88/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Tranquilla
Stack tecnologico
python, pytorch

Direzione di ricerca

Inizia da Matrix-Game-3/wan/distributed/ulysses.py:63-70 e confronta la keyword di attention() con la firma in wan/modules/attention.py:176. Esegui la riproduzione multi-GPU fornita oppure Matrix-Game-3/test.sh; il lavoro è completato quando il percorso documentato con --ulysses_size maggiore di 1 non solleva più un TypeError unexpected-keyword.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Summary

Every run with --ulysses_size > 1 crashes immediately in the first DiT block with a TypeError. This includes the configuration in the repo's own Matrix-Game-3/test.sh (which uses --ulysses_size 7/8), so the documented multi-GPU path does not run as shipped.

Tested at commit 71c3cd7f741311f8100f6cf9cde942b6c1378d11.

Error

File "Matrix-Game-3/wan/distributed/ulysses.py", line 63, in distributed_attention
    x = attention(
TypeError: attention() got an unexpected keyword argument 'fa_version'

Cause

wan/distributed/ulysses.py:63-70 calls attention() with fa_version=:

from ..modules.attention import attention
x = attention(
    q, k, v,
    k_lens=seq_lens,
    window_size=window_size,
    fa_version=fa_version,   # <-- wrong keyword
)

but the parameter in wan/modules/attention.py:176 is named version:

def attention(
    q, k, v,
    ...
    version=None,
):

Only this call site is affected. The other fa_version=fa_version call sites (sequence_parallel.py:307,364 and model.py:591,596,613,615,678,1015,1027) target functions that do declare fa_version, so they are correct.

Fix

One word, in wan/distributed/ulysses.py:69:

-        fa_version=fa_version,
+        version=fa_version,

Reproduce

torchrun --nproc_per_node=2 generate.py \
  --size 704*1280 --ckpt_dir Matrix-Game-3.0 \
  --image demo_images/001/image.png --prompt "..." \
  --ulysses_size 2 --dit_fsdp --t5_fsdp \
  --num_iterations 3 --num_inference_steps 3 --use_int8

Fails identically regardless of --fa_version, GPU model, or whether FlashAttention is installed — the TypeError is raised before any attention backend is selected.

Environment

2x RTX 5090 (sm_120), torch 2.11.0+cu128, Python 3.12, WSL2 Ubuntu 22.04. The bug is hardware-independent.

Related minor issues

Happy to open these separately if you'd prefer — noting them here since they came up while getting the model running on 2 GPUs:

  1. WAN_VAE_SEGMENT_SIZE is undocumented. inference_pipeline.py:643 reads it (default 4), and it decides whether inference fits in 32 GB. At 4 the VAE decode exhausts VRAM at 704x1280 even with --use_int8, --t5_cpu and --lightvae_pruning_rate 0.5; at 1 the same run completes. A README note or a --vae_segment_size flag would help.

  2. stream_decode swallows exceptions. wan/modules/vae2_2.py:1349 catches Exception, logs one line, and returns None. That None then surfaces ~200 lines away as AttributeError: 'NoneType' object has no attribute 'cpu' (inference_pipeline.py:652), hiding the real CUDA/OOM error. Adding traceback.format_exc() there would make debugging much easier.

Thanks for open-sourcing the weights and code — the one-word fix above got 2-GPU inference working end to end for us.

Lingua principale
Python
Stelle
2.3k
Fork
256
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di SkyworkAI/Matrix-Game

Tutte le issue di SkyworkAI/Matrix-Game

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.