[BUG] SIGSEGV in fused_attn_fwd when cu_seqlens_q != cu_seqlens_q_padded with FP8 blockwise
メンテナーはふだん 2 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 45/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
調査の方向性
添付された repro_fp8_padding_bug.py から開始し、FP8 fused-attention パスおよび ColumnParallelLinear の sequence-parallel サイズ計算を通じて DotProductAttention を追跡します。cu_seqlens_q と cu_seqlens_q_padded の使用箇所を比較し、Megatron-Core multi_latent_attention.py:567-574 も確認します。完了条件は、FP8 blockwise で segfault が発生しなくなり、出力がパディング済みの形状を維持し、attention masks が実トークンをマスクし、通信がパディング済みの境界を使用することです。
索引モデルが issue の本文から書いたものです。
説明
Environment
- TransformerEngine: 2.12.0
- Megatron-Core: 0.16.0
- PyTorch: 2.9.1
- CUDA: 12.9
- GPU: H100
Description
DotProductAttention with thd (packed sequence) format crashes with SIGSEGV when cu_seqlens_q differs from cu_seqlens_q_padded and FP8 blockwise recipe is used.
Per the TE documentation, cu_seqlens_q should contain actual token boundaries while cu_seqlens_q_padded contains the padded memory layout boundaries. This is needed when there is padding between sequences in a packed batch (e.g., for FP8 alignment).
Expected Behavior
Attention should:
- Use
cu_seqlens_qfor attention masking (only attend to real tokens) - Use
cu_seqlens_q_paddedfor memory layout (tensor indexing) - Return output tensor in padded layout (same shape as input Q)
Actual Behavior
- FP8 blockwise: SIGSEGV in
tex.fused_attn_fwdC++ kernel - FP8 delayed: Output tensor has unpadded size instead of padded, causing downstream shape mismatches
- BF16 (no FP8): Works correctly when cu_seqlens differ (non-FP8 backends handle it)
Impact
This bug makes FP8 training with sequence packing unusable when FP8 alignment padding is needed. The VERL framework (volcengine/verl) adds FP8 alignment padding for TE compatibility but passes identical cu_seqlens_q and cu_seqlens_q_padded as a workaround. This causes padding tokens to be visible to attention, corrupting the model output (training perplexity goes from 3.7 to 3055).
Reproduction
import torch
import transformer_engine.pytorch as te
from transformer_engine.pytorch.attention import DotProductAttention
# Setup
batch_size = 4
real_seqlens = [100, 150, 120, 130] # actual sequence lengths
padded_seqlens = [112, 160, 128, 144] # padded to 16-byte alignment for FP8
total_padded = sum(padded_seqlens)
hidden_dim = 1536
num_heads = 32
head_dim = hidden_dim // num_heads
# Create cu_seqlens
cu_seqlens_q = torch.tensor([0] + list(torch.cumsum(torch.tensor(real_seqlens), 0)),
dtype=torch.int32, device='cuda')
cu_seqlens_q_padded = torch.tensor([0] + list(torch.cumsum(torch.tensor(padded_seqlens), 0)),
dtype=torch.int32, device='cuda')
# Create Q, K, V in thd format (padded layout)
q = torch.randn(total_padded, num_heads, head_dim, device='cuda', dtype=torch.bfloat16)
k = torch.randn(total_padded, num_heads, head_dim, device='cuda', dtype=torch.bfloat16)
v = torch.randn(total_padded, num_heads, head_dim, device='cuda', dtype=torch.bfloat16)
# This works (cu_seqlens_q == cu_seqlens_q_padded):
attn = DotProductAttention(num_heads, head_dim, head_dim)
with te.fp8_autocast(enabled=True, fp8_recipe=te.recipe.BlockScaling()):
out = attn(q, k, v,
qkv_format='thd',
cu_seqlens_q=cu_seqlens_q_padded, # same as padded
cu_seqlens_kv=cu_seqlens_q_padded,
cu_seqlens_q_padded=cu_seqlens_q_padded,
cu_seqlens_kv_padded=cu_seqlens_q_padded,
max_seqlen_q=max(padded_seqlens),
max_seqlen_kv=max(padded_seqlens),
attn_mask_type='causal')
# This CRASHES (cu_seqlens_q != cu_seqlens_q_padded):
with te.fp8_autocast(enabled=True, fp8_recipe=te.recipe.BlockScaling()):
out = attn(q, k, v,
qkv_format='thd',
cu_seqlens_q=cu_seqlens_q, # actual boundaries
cu_seqlens_kv=cu_seqlens_q,
cu_seqlens_q_padded=cu_seqlens_q_padded, # padded boundaries
cu_seqlens_kv_padded=cu_seqlens_q_padded,
max_seqlen_q=max(real_seqlens),
max_seqlen_kv=max(real_seqlens),
attn_mask_type='causal')
# SIGSEGV ^^^
Core Problem: cu_seqlens_q used for BOTH attention masking AND SP communication
When cu_seqlens_q != cu_seqlens_q_padded:
- DotProductAttention: Works correctly in isolation (masks padding, returns padded-size output)
- ColumnParallelLinear with sequence_parallel=True: Uses
cu_seqlens_qfor allgather sizing → allgathered tensor gets UNPADDED size → breaks downstream RoPE/Linear that expect padded size
This means TE uses cu_seqlens_q for two conflicting purposes:
- Attention masking (should use actual/unpadded boundaries)
- SP allgather/scatter (should use padded boundaries)
Proposed fix: TE should separate these two uses. SP communication should ALWAYS use cu_seqlens_q_padded for sizing, while attention masking should use cu_seqlens_q.
Impact on Training
With VERL framework training DeepSeek 10B MoE with Megatron-Core:
- BF16 baseline: grad_norm=0.28, training_ppl=3.7
- FP8 E2E (cu_seqlens_q == cu_seqlens_q_padded): grad_norm=130-500, training_ppl=3055 (garbage)
- BF16 + FP8 padding only (no FP8 compute): grad_norm=1064, training_ppl=6.6 (proves padding is the cause)
Related Issues
- #1409 — NaN errors when cu_seqlens_q differs from cu_seqlens_q_padded
- #2391 — FA3 backend does not support
pad_between_seqs=True - #1929 — incorrect handling of cu_seqlens_q values with THD format
- #2186 — FusedAttention backward pass NaNs with THD+CP
- TE docs:
DotProductAttentiondocstring showscu_seqlens_q=[0,3,5,9]withcu_seqlens_q_padded=[0,4,8,13]as valid use case - Megatron-Core
multi_latent_attention.py:567-574expects cu_seqlens_q to differ from cu_seqlens_q_padded - Checked TE 2.13.0 release notes — none of these issues are fixed
- 主要言語
- Python
- スター
- 3.6k
- フォーク
- 851
- 平均マージ
- 5日 1時間
- マージ済み PR(30日)
- 52
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/TransformerEngine のほかの issue
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on itオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3647 ·
メンテナーはふだん 2 日以内に返信
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar対応中かも @sanjana658 が 1 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3636 · コメント 2 件 ·
メンテナーはふだん 2 日以内に返信
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it対応中かも @yuweih205 が 32 日前に担当しました。 オープンattention
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
NVIDIA/TransformerEngine#3481 · コメント 4 件 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
NVIDIA/TransformerEngine#2189 · コメント 7 件 · リアクション 5 件 ·
メンテナーはふだん 2 日以内に返信
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration builds対応中かも @ksivaman が今日担当しました。 オープン
難易度 4/5 3〜5日 初心者へのやさしさ 50/100
NVIDIA/TransformerEngine#3645 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 2 日以内に返信
NVIDIA/TransformerEngine の issue をすべて見る
似ている issue
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0オープンdocumentation priority: low size: S
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
メンテナーはふだん 1 日以内に返信
-
bug help wanted
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
メンテナーはふだん 1 日以内に返信
-
documentation
難易度 1/5 1時間未満 初心者へのやさしさ 65/100
ansys/pydpf-core#3547 ·
メンテナーはふだん 1 日以内に返信
-
core
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
vectorize-io/hindsight#5457 ·
メンテナーはふだん 1 日以内に返信
-
[Bug]: LangChain drops OpenAI Responses text blocks from session recording対応中かも @ktz03 が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
volcengine/OpenViking#5806 ·
メンテナーはふだん 1 日以内に返信