[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on it
Maintainers usually reply within 2 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 82/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- performance
Research direction
Read transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh, focusing on group_block_scaled_2d_tma_kernel and group_block_scaled_1d_tma_kernel at the two cited barrier sequences. Move the existing __syncthreads() before the leader's invalidation at both sites, then run the relevant grouped FP8 blockwise tests; done when all threads join before either barrier is invalidated.
Written by the indexing model from the issue text.
Description
Describe the bug
In transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh (main 5e99464), both TMA kernels (group_block_scaled_2d_tma_kernel and group_block_scaled_1d_tma_kernel) end their TMA load like this (lines 343–345, and the same at 630–632):
ptx::mbarrier_wait_parity(&tma_mbar, 0);
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);
__syncthreads();
Every thread polls tma_mbar, and thread 0 invalidates it before the block joins. Nothing makes the other threads finish polling first. Warps execute independently, so another warp may not have started its poll yet. Also, mbarrier.try_wait may return after a system-dependent time limit and be retried (§9.7.15.16.19), so a lagging thread can poll again after the invalidation. Lanes 1–31 of warp 0 are not covered either: the wait loop is per thread, so lanes can diverge.
PTX ISA 9.4 (§9.7.15.16.13, mbarrier.inval): "Performing any mbarrier operation except mbarrier.init on a memory location that does not contain a valid mbarrier object, results in undefined behaviour." The ISA's own examples put a barrier before @t0 mbarrier.inval.
A late poll on the invalidated word is undefined. If invalidation changes the word, that thread may never see the phase complete, and the block hangs at the __syncthreads(). The window is narrow: the other warps start polling within cycles, while the load (32 KB for a BF16/FP16 tile, 64 KB for FP32) takes, we estimate, microseconds. We have not observed a failure on hardware; this is a conformance fix.
Steps/Code to reproduce bug
Found by checking a reduced extract of this sequence (a 4 KB linear bulk copy instead of the tensor-map load), compiled with CUDA 12.9 for sm_90a, against the PTX ISA. It has not been reproduced on hardware. A targeted probe would delay one non-leader warp before its first poll, use bounded waits, and compare against a version that joins first.
Expected behavior
All threads have finished waiting on tma_mbar before it is invalidated. The fix is to move the existing barrier above the invalidation at both sites:
ptx::mbarrier_wait_parity(&tma_mbar, 0);
__syncthreads();
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);
TE already does this elsewhere, for example in cast/mxfp8/dequantize_mxfp8.cuh (sm_100 path) and fused_attn/flash_attn.cu, which reach a __syncthreads() before invalidating their barriers. tma_mbar is not used after the invalidation, so no other change is needed.
Environment overview
Source analysis at 5e99464; no runtime environment involved.
Device details
Hopper (sm_90a) code path.
Additional context
Reach: the grouped FP8 block-scaling quantize kernels on Hopper (2D, 1D columnwise and both, and 1D rowwise with dbias through bgrad_group_quantize). te.ops.GroupedLinear reaches them by default on Hopper with cuBLASLt 13.6 or later and a block-scaling recipe; the te.pytorch.GroupedLinear module only with the opt-in use_grouped_tensor=True (or the deprecated NVTE_GROUPED_LINEAR_USE_FUSED_GROUPED_GEMM=1).
- Dominant language
- Python
- Stars
- 3.6k
- Forks
- 851
- Avg merge
- 4d 14h
- Merged PRs (30d)
- 54
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/TransformerEngine
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarPossibly taken @sanjana658 claimed this 3 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA/TransformerEngine#3636 · 2 comments ·
Maintainers usually reply within 2 days
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itPossibly taken @yuweih205 claimed this 34 days ago. Openattention
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
NVIDIA/TransformerEngine#3481 · 4 comments ·
Maintainers usually reply within 2 days
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/TransformerEngine#2189 · 7 comments · 5 reactions ·
Maintainers usually reply within 2 days
-
Fused gemm + comm for CP A2A on BlackwellPossibly taken @cyanguwa claimed this 2 days ago. Open2.22 attention
NVIDIA/TransformerEngine#3664 · 1 assignee ·
Maintainers usually reply within 2 days
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration buildsPossibly taken @ksivaman claimed this 3 days ago. Open
Difficulty 4/5 3-5 days Newbie friendliness 50/100
NVIDIA/TransformerEngine#3645 · 1 comment · 1 assignee ·
Maintainers usually reply within 2 days
All issues in NVIDIA/TransformerEngine
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 83/100
PedestrianDynamics/pyFDS-Evac#766 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 1-3 hours Newbie friendliness 91/100
alchaincyf/nuwa-skill#86 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 2 days
-
Docs Needs Triage
Difficulty 1/5 Under an hour Newbie friendliness 88/100
pandas-dev/pandas#71055 ·
Maintainers usually reply within 1 day
-
[Bug]: graphify reads files that git's global ignore file hidesPossibly taken @smngvlkz claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Graphify-Labs/graphify#4335 · 1 comment ·
Maintainers usually reply within 1 day