[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on it
Maintainer thường phản hồi trong vòng 2 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 82/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- python
- Lĩnh vực
- performance
Hướng nghiên cứu
Đọc transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh, tập trung vào group_block_scaled_2d_tma_kernel và group_block_scaled_1d_tma_kernel tại hai chuỗi barrier được trích dẫn. Di chuyển __syncthreads() hiện có lên trước thao tác vô hiệu hóa của leader ở cả hai vị trí, sau đó chạy các bài kiểm thử FP8 blockwise nhóm liên quan; hoàn tất khi tất cả các luồng đều đồng bộ trước khi một trong hai barrier bị vô hiệu hóa.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the bug
In transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh (main 5e99464), both TMA kernels (group_block_scaled_2d_tma_kernel and group_block_scaled_1d_tma_kernel) end their TMA load like this (lines 343–345, and the same at 630–632):
ptx::mbarrier_wait_parity(&tma_mbar, 0);
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);
__syncthreads();
Every thread polls tma_mbar, and thread 0 invalidates it before the block joins. Nothing makes the other threads finish polling first. Warps execute independently, so another warp may not have started its poll yet. Also, mbarrier.try_wait may return after a system-dependent time limit and be retried (§9.7.15.16.19), so a lagging thread can poll again after the invalidation. Lanes 1–31 of warp 0 are not covered either: the wait loop is per thread, so lanes can diverge.
PTX ISA 9.4 (§9.7.15.16.13, mbarrier.inval): "Performing any mbarrier operation except mbarrier.init on a memory location that does not contain a valid mbarrier object, results in undefined behaviour." The ISA's own examples put a barrier before @t0 mbarrier.inval.
A late poll on the invalidated word is undefined. If invalidation changes the word, that thread may never see the phase complete, and the block hangs at the __syncthreads(). The window is narrow: the other warps start polling within cycles, while the load (32 KB for a BF16/FP16 tile, 64 KB for FP32) takes, we estimate, microseconds. We have not observed a failure on hardware; this is a conformance fix.
Steps/Code to reproduce bug
Found by checking a reduced extract of this sequence (a 4 KB linear bulk copy instead of the tensor-map load), compiled with CUDA 12.9 for sm_90a, against the PTX ISA. It has not been reproduced on hardware. A targeted probe would delay one non-leader warp before its first poll, use bounded waits, and compare against a version that joins first.
Expected behavior
All threads have finished waiting on tma_mbar before it is invalidated. The fix is to move the existing barrier above the invalidation at both sites:
ptx::mbarrier_wait_parity(&tma_mbar, 0);
__syncthreads();
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);
TE already does this elsewhere, for example in cast/mxfp8/dequantize_mxfp8.cuh (sm_100 path) and fused_attn/flash_attn.cu, which reach a __syncthreads() before invalidating their barriers. tma_mbar is not used after the invalidation, so no other change is needed.
Environment overview
Source analysis at 5e99464; no runtime environment involved.
Device details
Hopper (sm_90a) code path.
Additional context
Reach: the grouped FP8 block-scaling quantize kernels on Hopper (2D, 1D columnwise and both, and 1D rowwise with dbias through bgrad_group_quantize). te.ops.GroupedLinear reaches them by default on Hopper with cuBLASLt 13.6 or later and a block-scaling recipe; the te.pytorch.GroupedLinear module only with the opt-in use_grouped_tensor=True (or the deprecated NVTE_GROUPED_LINEAR_USE_FUSED_GROUPED_GEMM=1).
- Ngôn ngữ chính
- Python
- Star
- 3.6k
- Fork
- 851
- Merge trung bình
- 4 ngày 20 giờ
- Pull request đã merge (30 ngày)
- 56
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/TransformerEngine
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarCó thể đã có người làm @sanjana658 đã nhận 1 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA/TransformerEngine#3636 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itCó thể đã có người làm @yuweih205 đã nhận 32 ngày trước. Đang mởattention
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
NVIDIA/TransformerEngine#3481 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Increase MAX_TENSOR_NUMĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
NVIDIA/TransformerEngine#2189 · 7 bình luận · 5 reaction ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Fused gemm + comm for CP A2A on BlackwellCó thể đã có người làm @cyanguwa đã nhận hôm nay. Đang mở2.22 attention
NVIDIA/TransformerEngine#3664 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration buildsCó thể đã có người làm @ksivaman đã nhận 1 ngày trước. Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 50/100
NVIDIA/TransformerEngine#3645 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của NVIDIA/TransformerEngine
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
mishraprafful/multihull#150 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 66/100
python-caldav/caldav#735 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
mealie-recipes/mealie#8682 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
good first issue lane:repo
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100