Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[Bug] CUDA graph backward replay illegal memory access after validation with mixed NVFP4/MXFP8 Mamba

オープン
#3,316 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
静か
技術スタック
python, pytorch

調査の方向性

transformer_engine/pytorch/graph.py: backward から始め、記載されている GB200 構成を使って iteration-100 の validation から training への遷移を再現します。失敗する Mamba CUDA graph と混在した NVFP4/MXFP8/BF16 セットアップを、Mamba を削除するか E4M3 MXFP8 を使用する対照条件と比較します。完了条件は、validation 後に illegal-memory-access や segmentation fault を発生させず training が再開することです。

索引モデルが issue の本文から書いたものです。

説明

Describe the bug

Nemotron 3 Super training on GB200 consistently hits an illegal memory access around the iteration-100 validation boundary when using Transformer Engine CUDA graphs with a mixed per-module NVFP4, MXFP8, and BF16 recipe.

  • MXFP8 w/ cuda graphs for attn, mamba, moe_router, moe_preprocess: works
  • NVFP4 w/ cuda graphs for attn, mamba, moe_router, moe_preprocess: hits illegal memory access between iterations 90 and 100
  • NVFP4 w/ cuda graphs for attn, moe_router, moe_preprocess: works

The symptom initially appeared between the iteration 90 and 100 log points, but detailed timestamps show that iteration 100 and its 14 validation iterations both complete successfully. The failure occurs during the first training backward after validation. One rank reports a CUDA illegal-memory-access error from the Megatron Core backward call, while another rank can simultaneously segfault inside Transformer Engine CUDA graph backward replay. The failing ranks vary across executions.

Representative stacks:

train_step
  -> forward_backward_no_pipelining
  -> backward_step
  -> custom_backward
  -> Variable._execution_engine.run_backward
torch.AcceleratorError: CUDA error: an illegal memory access was encountered
transformer_engine/pytorch/graph.py: backward
  -> torch/cuda/graphs.py: replay
  -> at::cuda::CUDAGraph::replay
  -> cudaGraphLaunch
  -> cuGraphLaunch
Fatal Python error: Segmentation fault

Steps/Code to reproduce bug

Run Nemotron 3 Super training with this configuration:

Hardware: 16 GB200 nodes, 4 GPUs per node, 64 ranks
Parallelism: TP=2, PP=1, EP=64, ETP=1
Sequence length: 8192
Micro/global batch size: 1/512
Transformer implementation: transformer_engine
Attention backend: fused
CUDA graph implementation: transformer_engine
CUDA graph modules: [attn, mamba, moe_router, moe_preprocess]
CUDA graph warmup steps: 3
MoE dispatcher: flex with HybridEP, 32 SMs
Evaluation interval/iterations: 100/14
Manual GC interval: 100

The base quantization recipe is NVFP4 E2M1. The per-module precision configuration keeps attention QKV/projection, latent projections, and MTP in BF16, and uses MXFP8 for the Mamba output projection. The last 14 layers are also BF16.

The exact mixed-precision arguments used by the failing NVFP4 job were:

--bf16
--grad-reduce-in-bf16
--te-precision-config-file /mnt/artifacts/model/nemotron3_super_release_gb200/te_quant.cfg
--first-last-layers-bf16
--num-layers-at-start-in-bf16 0
--num-layers-at-end-in-bf16 14
--fp4-format e2m1
--fp4-recipe nvfp4

FP4 parameter gather was not enabled (fp4_param_gather=False).

The relevant runtime settings include:

CUDA_DEVICE_MAX_CONNECTIONS=32
NCCL_GRAPH_REGISTER=0
NCCL_NVLS_ENABLE=0
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
NUM_OF_HYBRID_EP_RANKS_PER_NVLINK_DOMAIN=64
NUM_OF_TOKENS_PER_CHUNK_COMBINE_API=128
NVTE_FWD_LAYERNORM_SM_MARGIN=20
NVTE_BWD_LAYERNORM_SM_MARGIN=20

Train through iteration 100, run validation, and resume training. In failing runs, iteration 100 completes, validation takes about 112 seconds and completes, then the first resumed training backward faults within several seconds.

What we tested and how we narrowed it down

  • The failure reproduced across multiple independent executions and retries on different nodes and ranks.
  • It reproduced with Transformer Engine 2.18 at two revisions, including 2.18.0+e0a871e6 with the merged weight-preswizzling stale-pointer fix and the later single-grouped-weight fixes present. Those fixes did not resolve this end-to-end failure.
  • It reproduced with both NVSHMEM 3.6.5 and 3.4.5, so the NVSHMEM upgrade is not the cause.
  • We aligned CUDA, NCCL, allocator, layernorm-margin, HybridEP-domain, and token-chunk settings with known working GB200 recipes. The failure remained.
  • A Transformer Engine 2.17 plus cuDNN 9.21 diagnostic completed only five iterations. It did not exercise the iteration-100 boundary, so it is not evidence for or against a version regression.
  • Keeping the updated configuration but removing only Mamba from the CUDA graph module list allowed validation at iteration 100 to complete and training to continue cleanly through iteration 120, with zero skipped or NaN iterations.
  • Keeping Mamba in the full CUDA graph module list but replacing the mixed per-module NVFP4/MXFP8/BF16 recipe with an E4M3 MXFP8 recipe plus BF16 boundary layers allowed validation at iteration 100 to complete and training to continue cleanly through iteration 170, again with zero skipped or NaN iterations.

These controls narrow the problem to an interaction that requires Mamba CUDA graph capture and the mixed per-module NVFP4/MXFP8/BF16 recipe, or state introduced by that recipe across the train-to-validation-to-train transition. It does not look like a generic HybridEP, NVSHMEM, or Mamba training failure.

Expected behavior

Validation should not leave stale or incompatible quantization or CUDA graph state. Training should resume after validation without an illegal memory access during backward graph replay.

Environment overview

  • Environment location: Docker on an internal bare-metal GB200 cluster
  • Transformer Engine install: built from source from the Megatron Core dependency lock during the Docker image build
  • Base image: NVIDIA PyTorch 26.06-derived image

Environment details

  • Python: 3.12
  • PyTorch: 2.13.0a0+8145d630e8.nv26.6
  • Megatron Core: 0.19.0+af9e4408d
  • Transformer Engine: 2.18.0+e0a871e6
  • CUDA: 13.3
  • NCCL: 2.30.7+cuda13.3

Device details

  • GPU model: NVIDIA GB200
  • Topology: 16 nodes, 4 GPUs per node

Additional context

The CUDA error is asynchronous, so the Python error site alone does not identify the captured kernel that first corrupts memory. The simultaneous native stack on another rank localizes the visible failure to Transformer Engine backward graph replay, but we have not yet isolated the exact Mamba or quantization kernel. We can run a launch-blocking diagnostic or help reduce this to a smaller reproducer if there is a preferred instrumentation path.

主要言語
Python
スター
3.6k
フォーク
851
平均マージ
4日 20時間
マージ済み PR(30日)
56

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/TransformerEngine のほかの issue

NVIDIA/TransformerEngine の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。