Performance bottleneck in TransformerLayer and Attention when using attention mask
メンテナーはふだん 2 日以内に返信
@cyanguwa がすでに取り組んでいます。
2026年3月2日 から。
評価
この issue はまだ評価されていません。
説明
Describe the bug
Dear Nvidia experts,
When attention mask is used, these three lines will cause very significant slow down since they are looping over all the items in the batch.
Performance bottleneck in TransformerLayer and Attention.
assertion the type of attention mask
similarly another assertion the type of attention mask
See below for profiling trace:
As you can see from the largest bar from the bottom rows, operations taking most of the time in each transformer is now transformer_engine/pytorch/attention/dot_product_attention/utils.py(1518): get_indices, Similarly the other two assertion is also cause significant slow down.
Steps/Code to reproduce bug
This problem should manifest in any profiling run with data that have attention mask in the input of forward()
Expected behavior
Multiple times of slow down when calling forward of TransformerLayer.
- 主要言語
- Python
- スター
- 3.6k
- フォーク
- 844
- 平均マージ
- 4日 55分
- マージ済み PR(30日)
- 51
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/TransformerEngine のほかの issue
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3636 ·
メンテナーはふだん 2 日以内に返信
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it対応中かも @yuweih205 が 29 日前に担当しました。 オープンattention
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
NVIDIA/TransformerEngine#3481 · コメント 4 件 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
NVIDIA/TransformerEngine#2189 · コメント 7 件 · リアクション 5 件 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 54/100
NVIDIA/TransformerEngine#3640 · コメント 5 件 ·
メンテナーはふだん 2 日以内に返信
-
[BUG] Grouped MXFP8 quantization is not concurrency safe with multiple streams対応中かも @kainzhong が 1 日前に担当しました。 オープンbug
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
NVIDIA/TransformerEngine#3630 ·
メンテナーはふだん 2 日以内に返信
NVIDIA/TransformerEngine の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 3 日以内に返信
-
Negation with "not" and "no" is ignored during sentiment analysis対応中かも @vivek-3728 が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
techcsispit/mess-mood#11 · コメント 1 件 ·
-
changelog investigate
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
ramnes/notion-sdk-py#408 ·
-
good first issue
難易度 2/5 1〜3時間 初心者へのやさしさ 83/100
btclib-org/btclib-wallet#267 ·
メンテナーはふだん 1 日以内に返信
-
good first issue tech-debt
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
knnmelprop/YAADO#111 ·