Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Performance bottleneck in TransformerLayer and Attention when using attention mask

オープン
#2,703 コメント 1 件 リアクション 0 件 担当者 1 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

@cyanguwa がすでに取り組んでいます。

2026年3月2日 から。

評価

この issue はまだ評価されていません。

説明

bug

Describe the bug

Dear Nvidia experts,

When attention mask is used, these three lines will cause very significant slow down since they are looping over all the items in the batch.

Performance bottleneck in TransformerLayer and Attention.

function get_indices

assertion the type of attention mask

similarly another assertion the type of attention mask

See below for profiling trace:

Image

As you can see from the largest bar from the bottom rows, operations taking most of the time in each transformer is now transformer_engine/pytorch/attention/dot_product_attention/utils.py(1518): get_indices, Similarly the other two assertion is also cause significant slow down.

Steps/Code to reproduce bug

This problem should manifest in any profiling run with data that have attention mask in the input of forward()

Expected behavior

Multiple times of slow down when calling forward of TransformerLayer.

主要言語
Python
スター
3.6k
フォーク
844
平均マージ
4日 55分
マージ済み PR(30日)
51

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/TransformerEngine のほかの issue

NVIDIA/TransformerEngine の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。