Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Performance bottleneck in TransformerLayer and Attention when using attention mask

未关闭
#2,703 1 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

维护者通常 2 天内回复

@cyanguwa 已经在做这个了。

开始于 2026年3月2日。

评估

这个 Issue 还没有评估数据。

描述

bug

Describe the bug

Dear Nvidia experts,

When attention mask is used, these three lines will cause very significant slow down since they are looping over all the items in the batch.

Performance bottleneck in TransformerLayer and Attention.

function get_indices

assertion the type of attention mask

similarly another assertion the type of attention mask

See below for profiling trace:

Image

As you can see from the largest bar from the bottom rows, operations taking most of the time in each transformer is now transformer_engine/pytorch/attention/dot_product_attention/utils.py(1518): get_indices, Similarly the other two assertion is also cause significant slow down.

Steps/Code to reproduce bug

This problem should manifest in any profiling run with data that have attention mask in the input of forward()

Expected behavior

Multiple times of slow down when calling forward of TransformerLayer.

主要语言
Python
星标
3.6k
派生
844
平均合并
4 天 55 分钟
30 天内合并 PR
51

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

NVIDIA/TransformerEngine 的其他 Issue

查看 NVIDIA/TransformerEngine 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。