Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Performance bottleneck in TransformerLayer and Attention when using attention mask

Aperta
#2,703 1 commento 0 reazioni 1 assegnatario Vedi su GitHub

I maintainer di solito rispondono entro 2 giorni

@cyanguwa ci sta già lavorando.

Dal 2/3/2026.

Valutazione

Questa issue non è ancora stata valutata.

Descrizione

bug

Describe the bug

Dear Nvidia experts,

When attention mask is used, these three lines will cause very significant slow down since they are looping over all the items in the batch.

Performance bottleneck in TransformerLayer and Attention.

function get_indices

assertion the type of attention mask

similarly another assertion the type of attention mask

See below for profiling trace:

Image

As you can see from the largest bar from the bottom rows, operations taking most of the time in each transformer is now transformer_engine/pytorch/attention/dot_product_attention/utils.py(1518): get_indices, Similarly the other two assertion is also cause significant slow down.

Steps/Code to reproduce bug

This problem should manifest in any profiling run with data that have attention mask in the input of forward()

Expected behavior

Multiple times of slow down when calling forward of TransformerLayer.

Lingua principale
Python
Stelle
3.6k
Fork
851
Merge medio
4g 15h
PR unite (30g)
51

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di NVIDIA/TransformerEngine

Tutte le issue di NVIDIA/TransformerEngine

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.