[PyTorch][Attention] THD P2P context-parallel regression when padded cu_seqlens are value-equal but not object-identical
メンテナーはふだん 2 日以内に返信
まだ誰も着手していません。
- #3506 @RPalmr による — マージされずにクローズ
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 38/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
調査の方向性
4cd705b75394563c0246bdddfa5d3148106c9285 のコミットで導入された PyTorch DotProductAttention の padding 検出器から着手し、get_cu_seqlens_on_cp_rank と THD の dQ/dK/dV の末尾をゼロ化するパスを調査します。提供されたベンチマークを使って pad_between_seqs を交互に指定するケースを再現し、カーネルの処理量と実行時間を比較します。値が同等なメタデータによって不要な padding 処理を回避しつつ、数値的な動作とグラフセーフな実行が維持されれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Describe the bug
Commit 4cd705b75394563c0246bdddfa5d3148106c9285 introduces a PyTorch DotProductAttention performance regression under this combination:
qkv_format="thd"- P2P context parallelism
pad_between_seqs=None(automatic detection)- both regular and padded cumulative sequence-length tensors are provided
- the padded and unpadded tensors are distinct objects but have identical relevant values, so there is no actual padding between sequences
This is observable during ordinary eager training; CUDA graph capture does not need to be enabled.
The new automatic detection uses tensor object identity as a proxy for padding semantics:
if cu_seqlens_q_padded is cu_seqlens_q:
pad_between_seqs = False
elif cu_seqlens_q_padded is not None or cu_seqlens_kv_padded is not None:
pad_between_seqs = True
Thus independently allocated but value-identical tensors are classified as pad_between_seqs=True. With P2P context parallelism, that selects the path that repeatedly calls get_cu_seqlens_on_cp_rank, rather than the cheaper no-inter-sequence-padding path.
The same commit also adds THD dQ/dK/dV tail-zeroing operations. They launch arange/compare/masked-fill work even when the valid endpoint already equals the tensor endpoint and there is no tail to clear.
Steps/Code to reproduce bug
-
Create a BF16
DotProductAttentionmodule withqkv_format="thd"and a four-rank P2P context-parallel group. -
Provide independently allocated cumulative sequence-length tensors with identical values:
cu_seqlens = torch.tensor([0, sequence_length], dtype=torch.int32, device="cuda") cu_seqlens_padded = cu_seqlens.clone() assert cu_seqlens_padded is not cu_seqlens assert torch.equal(cu_seqlens_padded, cu_seqlens) -
Run repeated attention forward/backward calls, alternating these two cases in the same process:
- automatic detection:
pad_between_seqs=None - known-correct metadata:
pad_between_seqs=False
- automatic detection:
-
Discard warmup and compare steady-state timings. A single attention forward/backward call shows a small direct overhead. The impact becomes much larger in an attention-heavy training schedule where the branch is exercised repeatedly and interacts with context-parallel stream scheduling.
We also performed a controlled source-level reverse experiment on Transformer Engine 2.18.0+27486e03. All arms used the same process, allocation, inputs, configuration, and byte-identical compiled Transformer Engine extensions; only the Python attention hunks from the cited commit differed. Each arm used 50 post-warmup iterations.
| Variant | Mean iteration time | Median | Delta vs. stock |
|---|---|---|---|
| Stock | 595.084 ms | 589.000 ms | — |
| Revert padding detector only | 560.958 ms | 556.800 ms | -5.735% |
| Revert gradient zero-fill only | 577.764 ms | 574.950 ms | -2.911% |
| Revert both | 554.254 ms | 549.650 ms | -6.861% |
An ABBA repetition of stock and the full reverse patch measured a 6.276% aggregate iteration-time improvement with the reverse patch. The stock and reverse-patched order drift was 1.534% and 0.861%, respectively.
Across all ranks in two Nsight Systems trials, the full reverse patch reduced the five-step trace span by 7.411% on average, with every paired rank faster. Over five steps it removed, per rank:
- 2,400 helper-generated kernel launches associated with
get_cu_seqlens_on_cp_rank - 900 masked-fill kernels from the new gradient tail-zeroing blocks
The source reversal reduced main-stream kernel work by 31.702 ms/rank and main-stream gaps by 237.154 ms/rank over the captured five-step window. These are separate trace observations, not additive wall-time attribution. Numerical-health checks remained clean.
Expected behavior
When the padded and unpadded cumulative sequence-length tensors have equal relevant values, automatic detection should not select the inter-sequence-padding path solely because they are different Python objects.
Could the API carry graph-safe padding metadata explicitly, or otherwise avoid using object identity as the semantic proxy? The tail-zeroing work could also be gated when metadata establishes that no gradient tail exists, while preserving CUDA graph compatibility.
The current workaround for callers that know there is no inter-sequence padding is to pass pad_between_seqs=False explicitly.
Environment overview
- Environment location: containerized bare-metal system
- Transformer Engine:
2.18.0+27486e03 - Installation: preinstalled container package
Environment details
- Python: 3.12.3
- PyTorch:
2.13.0a0+8145d630e8.nv26.6.54250401 - CUDA reported by PyTorch: 13.3
- cuDNN: compiled against 9.23; node-visible runtime 9.21.1
The causal comparison used one unchanged environment, so the cuDNN packaging detail was identical across all variants.
Device details
- 4x NVIDIA H100 80GB HBM3
Additional context
The detector is the primary contributor. After removing the zero-fill blocks, reverting the detector still improved iteration time by 4.069%. Once the detector was corrected, removing zero-fill added another 1.195%. The effects overlap on the same context-parallel critical path and therefore should not be added independently.
- 主要言語
- Python
- スター
- 3.6k
- フォーク
- 851
- 平均マージ
- 5日 1時間
- マージ済み PR(30日)
- 52
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/TransformerEngine のほかの issue
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on itオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3647 ·
メンテナーはふだん 2 日以内に返信
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar対応中かも @sanjana658 が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3636 · コメント 2 件 ·
メンテナーはふだん 2 日以内に返信
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it対応中かも @yuweih205 が 31 日前に担当しました。 オープンattention
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
NVIDIA/TransformerEngine#3481 · コメント 4 件 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
NVIDIA/TransformerEngine#2189 · コメント 7 件 · リアクション 5 件 ·
メンテナーはふだん 2 日以内に返信
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration builds対応中かも @ksivaman が今日担当しました。 オープン
難易度 4/5 3〜5日 初心者へのやさしさ 50/100
NVIDIA/TransformerEngine#3645 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 2 日以内に返信
NVIDIA/TransformerEngine の issue をすべて見る
似ている issue
-
first
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
AcademySoftwareFoundation/rmtc#54 · コメント 1 件 ·
-
feature/cohorts feature/feature-flags team/feature-flags
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
メンテナーはふだん 1 日以内に返信
-
License examples/ as MIT対応中かも @PGrayCS が今日担当しました。 オープンdocumentation enhancement example good first issue
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
speedyk-005/yasbd-lib#383 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
interactions-py/interactions.py#1827 ·
-
Managed start can fail when OpenVMM reads its control capability before NVX writes it対応中かも @ppenna が今日担当しました。 オープンbug
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
メンテナーはふだん 1 日以内に返信