Performance bottleneck in TransformerLayer and Attention when using attention mask
I maintainer di solito rispondono entro 2 giorni
@cyanguwa ci sta già lavorando.
Dal 2/3/2026.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Describe the bug
Dear Nvidia experts,
When attention mask is used, these three lines will cause very significant slow down since they are looping over all the items in the batch.
Performance bottleneck in TransformerLayer and Attention.
assertion the type of attention mask
similarly another assertion the type of attention mask
See below for profiling trace:
As you can see from the largest bar from the bottom rows, operations taking most of the time in each transformer is now transformer_engine/pytorch/attention/dot_product_attention/utils.py(1518): get_indices, Similarly the other two assertion is also cause significant slow down.
Steps/Code to reproduce bug
This problem should manifest in any profiling run with data that have attention mask in the input of forward()
Expected behavior
Multiple times of slow down when calling forward of TransformerLayer.
- Lingua principale
- Python
- Stelle
- 3.6k
- Fork
- 851
- Merge medio
- 4g 15h
- PR unite (30g)
- 51
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di NVIDIA/TransformerEngine
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
NVIDIA/TransformerEngine#3636 ·
I maintainer di solito rispondono entro 2 giorni
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itForse già presa @yuweih205 l’ha presa 31 giorni fa. Apertaattention
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
NVIDIA/TransformerEngine#3481 · 4 commenti ·
I maintainer di solito rispondono entro 2 giorni
-
Increase MAX_TENSOR_NUMApertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
NVIDIA/TransformerEngine#2189 · 7 commenti · 5 reazioni ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 50/100
NVIDIA/TransformerEngine#3645 ·
I maintainer di solito rispondono entro 2 giorni
-
enhancement
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
NVIDIA/TransformerEngine#3644 ·
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di NVIDIA/TransformerEngine
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
UKGovernmentBEIS/inspect_ai#5781 ·
I maintainer di solito rispondono entro 2 giorni
-
Bump .cicd to wamp-cicd 4c2f9ac: `just land` refuses open A18 decisions, `just where` lists themAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 82/100
crossbario/cfxdb#139 ·
-
Bump .cicd to wamp-cicd 4c2f9ac: `just land` refuses open A18 decisions, `just where` lists themAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 84/100
crossbario/txaio#241 ·
-
UX
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
mediajunkie/piper-morgan-product#1963 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 64/100