[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration builds
Maintainer thường phản hồi trong vòng 2 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 50/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Lĩnh vực
- performance
Hướng nghiên cứu
Start with transformer_engine/pytorch/graph.py, especially the registration loop around the linked line 484, and compare its capability check with the PyTorch implementations cited in the issue. Then trace the graph setup and existing CUDA graph RNG tests; add coverage for automatic versus required registration, multiple RNG states, and advancement across replay. Done means preserving correct capture/replay while avoiding deprecated no-op calls and their warning flood.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the bug
TE's CUDA graph setup calls CUDAGraph.register_generator_state() once per active RNG tracker state for every forward, backward, and wgrad graph. On the affected NVIDIA PyTorch 26.08 build, this API is a deprecated no-op that emits a C++ warning on every call because registration now happens automatically.
This produces a warning flood during distributed training. With pipeline parallelism, Megatron Core supplies capture samples per layer and microbatch, multiplying the calls. The investigation in NVIDIA-NeMo/Megatron-Bridge#6336 reports approximately 8K–26K warnings per rank for 256-GPU DeepSeek V3 and Qwen3 235B bf16 runs. Some ranks stall writing stderr while others finish capture and wait in collectives; the jobs then hit the 600-second NCCL watchdog timeout after their CUDA graph warmup iterations.
Steps/Code to reproduce bug
The affected path is:
Bridge training → MCore TECudaGraphHelper.create_cudagraphs() → TE make_graphed_callables() → _make_graphed_callables() → register_generator_state().
The registration loop in the Bridge-pinned TE revision registers all three graph types whenever graph_safe_rng_available() is true. That capability check only checks method availability; it does not distinguish required registration from a deprecated no-op.
On the affected PyTorch build, the underlying warning can be demonstrated independently of a full training run:
import torch
generator = torch.Generator(device="cuda")
graph = torch.cuda.CUDAGraph()
for _ in range(3):
graph.register_generator_state(generator)
This is a minimal illustration of the repeated-warning trigger, not a standalone reproduction of the distributed timeout. The full affected workload uses TE graphs, active RNG tracker states, and pipeline-parallel layer/microbatch capture; the four recipe names and training context are listed in NVIDIA-NeMo/Megatron-Bridge#6336.
The PyTorch implementation at 4fdf77b940 confirms the per-call TORCH_WARN_DEPRECATION and no-op behavior.
Expected behavior / requested TE fix
Please make TE's internal RNG registration conditional on the installed PyTorch implementation's requirements:
- Skip explicit registration when PyTorch handles it automatically.
- Preserve explicit registration on builds where it is required for correct capture/replay.
- Cover both behaviors with tests, including multiple RNG states and RNG advancement across replay, without requiring callers to suppress unrelated C++ warnings.
This should be handled centrally in TE's graph setup so all callers benefit. The inspected make_graphed_callables() API has no registration-policy argument through which MCore could control the inner loop.
Environment overview / device details
- Affected environment: NVIDIA PyTorch 26.08-based training container; reported torch build
2.14.0a0+4fdf77b940.nv26.08. - Workloads: 256-GPU DeepSeek V3 and Qwen3 235B-A22B bf16 training on GB200/GB300 with TE CUDA graphs.
- Bridge revision for the workaround:
93915172fa1711cca9d3a7fa0721fb15114cd8e8. - Bridge's TE source pin:
131157b5750b8578cb8efb78554f6fe6a671cc53(2.20.0source version). - The same registration loop is also present in inspected TE v2.18 (
27486e03), release_v2.20 (6ea2a74a), and main (5e994640,2.21.0.dev0). The latter revisions were source-inspected, not independently reproduced on GPUs for this report.
Additional context: workaround and compatibility
NVIDIA-NeMo/Megatron-Bridge#6336 sets TORCH_CPP_LOG_LEVEL=ERROR in four affected recipes before the training interpreter starts. This is a workaround: it suppresses the flood but leaves the unnecessary calls in place, hides unrelated C++ warnings, and does not cover other TE graph callers. We need a TE fix so the global warning suppression can be removed.
Please do not use a blanket torch >= 2.14 skip. NVIDIA/Megatron-LM#7725 reports that the NVIDIA-packaged NGC 26.09 build still requires explicit registration and fails capture if it is skipped, even though the upstream base source also contains the deprecated no-op. Both builds advertise PyTorch 2.14 prereleases. That PR proposes a conservative exact-build exception for the verified 26.08 build; the vendor patch responsible for the packaged-source discrepancy remains unconfirmed. A reliable capability contract would be preferable to a version-only decision.
- Ngôn ngữ chính
- Python
- Star
- 3.6k
- Fork
- 851
- Merge trung bình
- 4 ngày 15 giờ
- Pull request đã merge (30 ngày)
- 51
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/TransformerEngine
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA/TransformerEngine#3636 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itCó thể đã có người làm @yuweih205 đã nhận 31 ngày trước. Đang mởattention
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
NVIDIA/TransformerEngine#3481 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Increase MAX_TENSOR_NUMĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
NVIDIA/TransformerEngine#2189 · 7 bình luận · 5 reaction ·
Maintainer thường phản hồi trong vòng 2 ngày
-
enhancement
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
NVIDIA/TransformerEngine#3644 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 54/100
NVIDIA/TransformerEngine#3640 · 8 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của NVIDIA/TransformerEngine
Issue tương tự
-
first
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
AcademySoftwareFoundation/rmtc#54 · 1 bình luận ·
-
feature/cohorts feature/feature-flags team/feature-flags
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
Maintainer thường phản hồi trong vòng 1 ngày
-
License examples/ as MITCó thể đã có người làm @PGrayCS đã nhận hôm nay. Đang mởdocumentation enhancement example good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
speedyk-005/yasbd-lib#383 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
interactions-py/interactions.py#1827 ·
-
Managed start can fail when OpenVMM reads its control capability before NVX writes itCó thể đã có người làm @ppenna đã nhận hôm nay. Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày