[Bug] CUDA graph backward replay illegal memory access after validation with mixed NVFP4/MXFP8 Mamba
メンテナーはふだん 2 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
調査の方向性
transformer_engine/pytorch/graph.py: backward から始め、記載されている GB200 構成を使って iteration-100 の validation から training への遷移を再現します。失敗する Mamba CUDA graph と混在した NVFP4/MXFP8/BF16 セットアップを、Mamba を削除するか E4M3 MXFP8 を使用する対照条件と比較します。完了条件は、validation 後に illegal-memory-access や segmentation fault を発生させず training が再開することです。
索引モデルが issue の本文から書いたものです。
説明
Describe the bug
Nemotron 3 Super training on GB200 consistently hits an illegal memory access around the iteration-100 validation boundary when using Transformer Engine CUDA graphs with a mixed per-module NVFP4, MXFP8, and BF16 recipe.
- MXFP8 w/ cuda graphs for attn, mamba, moe_router, moe_preprocess: works
- NVFP4 w/ cuda graphs for attn, mamba, moe_router, moe_preprocess: hits illegal memory access between iterations 90 and 100
- NVFP4 w/ cuda graphs for attn, moe_router, moe_preprocess: works
The symptom initially appeared between the iteration 90 and 100 log points, but detailed timestamps show that iteration 100 and its 14 validation iterations both complete successfully. The failure occurs during the first training backward after validation. One rank reports a CUDA illegal-memory-access error from the Megatron Core backward call, while another rank can simultaneously segfault inside Transformer Engine CUDA graph backward replay. The failing ranks vary across executions.
Representative stacks:
train_step
-> forward_backward_no_pipelining
-> backward_step
-> custom_backward
-> Variable._execution_engine.run_backward
torch.AcceleratorError: CUDA error: an illegal memory access was encountered
transformer_engine/pytorch/graph.py: backward
-> torch/cuda/graphs.py: replay
-> at::cuda::CUDAGraph::replay
-> cudaGraphLaunch
-> cuGraphLaunch
Fatal Python error: Segmentation fault
Steps/Code to reproduce bug
Run Nemotron 3 Super training with this configuration:
Hardware: 16 GB200 nodes, 4 GPUs per node, 64 ranks
Parallelism: TP=2, PP=1, EP=64, ETP=1
Sequence length: 8192
Micro/global batch size: 1/512
Transformer implementation: transformer_engine
Attention backend: fused
CUDA graph implementation: transformer_engine
CUDA graph modules: [attn, mamba, moe_router, moe_preprocess]
CUDA graph warmup steps: 3
MoE dispatcher: flex with HybridEP, 32 SMs
Evaluation interval/iterations: 100/14
Manual GC interval: 100
The base quantization recipe is NVFP4 E2M1. The per-module precision configuration keeps attention QKV/projection, latent projections, and MTP in BF16, and uses MXFP8 for the Mamba output projection. The last 14 layers are also BF16.
The exact mixed-precision arguments used by the failing NVFP4 job were:
--bf16
--grad-reduce-in-bf16
--te-precision-config-file /mnt/artifacts/model/nemotron3_super_release_gb200/te_quant.cfg
--first-last-layers-bf16
--num-layers-at-start-in-bf16 0
--num-layers-at-end-in-bf16 14
--fp4-format e2m1
--fp4-recipe nvfp4
FP4 parameter gather was not enabled (fp4_param_gather=False).
The relevant runtime settings include:
CUDA_DEVICE_MAX_CONNECTIONS=32
NCCL_GRAPH_REGISTER=0
NCCL_NVLS_ENABLE=0
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
NUM_OF_HYBRID_EP_RANKS_PER_NVLINK_DOMAIN=64
NUM_OF_TOKENS_PER_CHUNK_COMBINE_API=128
NVTE_FWD_LAYERNORM_SM_MARGIN=20
NVTE_BWD_LAYERNORM_SM_MARGIN=20
Train through iteration 100, run validation, and resume training. In failing runs, iteration 100 completes, validation takes about 112 seconds and completes, then the first resumed training backward faults within several seconds.
What we tested and how we narrowed it down
- The failure reproduced across multiple independent executions and retries on different nodes and ranks.
- It reproduced with Transformer Engine 2.18 at two revisions, including
2.18.0+e0a871e6with the merged weight-preswizzling stale-pointer fix and the later single-grouped-weight fixes present. Those fixes did not resolve this end-to-end failure. - It reproduced with both NVSHMEM 3.6.5 and 3.4.5, so the NVSHMEM upgrade is not the cause.
- We aligned CUDA, NCCL, allocator, layernorm-margin, HybridEP-domain, and token-chunk settings with known working GB200 recipes. The failure remained.
- A Transformer Engine 2.17 plus cuDNN 9.21 diagnostic completed only five iterations. It did not exercise the iteration-100 boundary, so it is not evidence for or against a version regression.
- Keeping the updated configuration but removing only Mamba from the CUDA graph module list allowed validation at iteration 100 to complete and training to continue cleanly through iteration 120, with zero skipped or NaN iterations.
- Keeping Mamba in the full CUDA graph module list but replacing the mixed per-module NVFP4/MXFP8/BF16 recipe with an E4M3 MXFP8 recipe plus BF16 boundary layers allowed validation at iteration 100 to complete and training to continue cleanly through iteration 170, again with zero skipped or NaN iterations.
These controls narrow the problem to an interaction that requires Mamba CUDA graph capture and the mixed per-module NVFP4/MXFP8/BF16 recipe, or state introduced by that recipe across the train-to-validation-to-train transition. It does not look like a generic HybridEP, NVSHMEM, or Mamba training failure.
Expected behavior
Validation should not leave stale or incompatible quantization or CUDA graph state. Training should resume after validation without an illegal memory access during backward graph replay.
Environment overview
- Environment location: Docker on an internal bare-metal GB200 cluster
- Transformer Engine install: built from source from the Megatron Core dependency lock during the Docker image build
- Base image: NVIDIA PyTorch 26.06-derived image
Environment details
- Python: 3.12
- PyTorch:
2.13.0a0+8145d630e8.nv26.6 - Megatron Core:
0.19.0+af9e4408d - Transformer Engine:
2.18.0+e0a871e6 - CUDA: 13.3
- NCCL:
2.30.7+cuda13.3
Device details
- GPU model: NVIDIA GB200
- Topology: 16 nodes, 4 GPUs per node
Additional context
The CUDA error is asynchronous, so the Python error site alone does not identify the captured kernel that first corrupts memory. The simultaneous native stack on another rank localizes the visible failure to Transformer Engine backward graph replay, but we have not yet isolated the exact Mamba or quantization kernel. We can run a launch-blocking diagnostic or help reduce this to a smaller reproducer if there is a preferred instrumentation path.
- 主要言語
- Python
- スター
- 3.6k
- フォーク
- 851
- 平均マージ
- 4日 20時間
- マージ済み PR(30日)
- 56
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/TransformerEngine のほかの issue
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on itオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3647 ·
メンテナーはふだん 2 日以内に返信
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar対応中かも @sanjana658 が 1 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3636 · コメント 2 件 ·
メンテナーはふだん 2 日以内に返信
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it対応中かも @yuweih205 が 32 日前に担当しました。 オープンattention
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
NVIDIA/TransformerEngine#3481 · コメント 4 件 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
NVIDIA/TransformerEngine#2189 · コメント 7 件 · リアクション 5 件 ·
メンテナーはふだん 2 日以内に返信
-
Fused gemm + comm for CP A2A on Blackwell対応中かも @cyanguwa が今日担当しました。 オープン2.22 attention
NVIDIA/TransformerEngine#3664 · 担当者 1 名 ·
メンテナーはふだん 2 日以内に返信
NVIDIA/TransformerEngine の issue をすべて見る
似ている issue
-
enhancement good first issue Stellar Wave trivial
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
StellarCanary/ProtocolCanary-Fixtures#258 ·
メンテナーはふだん 1 日以内に返信
-
github_actions
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
Hochfrequenz/aibap.mcp#578 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
mishraprafful/multihull#150 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
python-caldav/caldav#735 ·
メンテナーはふだん 1 日以内に返信