Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on it

Open Beginner friendly
#3,647 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 2 days

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
82/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
performance

Research direction

Read transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh, focusing on group_block_scaled_2d_tma_kernel and group_block_scaled_1d_tma_kernel at the two cited barrier sequences. Move the existing __syncthreads() before the leader's invalidation at both sites, then run the relevant grouped FP8 blockwise tests; done when all threads join before either barrier is invalidated.

Written by the indexing model from the issue text.

Description

Describe the bug

In transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh (main 5e99464), both TMA kernels (group_block_scaled_2d_tma_kernel and group_block_scaled_1d_tma_kernel) end their TMA load like this (lines 343–345, and the same at 630–632):

ptx::mbarrier_wait_parity(&tma_mbar, 0);
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);
__syncthreads();

Every thread polls tma_mbar, and thread 0 invalidates it before the block joins. Nothing makes the other threads finish polling first. Warps execute independently, so another warp may not have started its poll yet. Also, mbarrier.try_wait may return after a system-dependent time limit and be retried (§9.7.15.16.19), so a lagging thread can poll again after the invalidation. Lanes 1–31 of warp 0 are not covered either: the wait loop is per thread, so lanes can diverge.

PTX ISA 9.4 (§9.7.15.16.13, mbarrier.inval): "Performing any mbarrier operation except mbarrier.init on a memory location that does not contain a valid mbarrier object, results in undefined behaviour." The ISA's own examples put a barrier before @t0 mbarrier.inval.

A late poll on the invalidated word is undefined. If invalidation changes the word, that thread may never see the phase complete, and the block hangs at the __syncthreads(). The window is narrow: the other warps start polling within cycles, while the load (32 KB for a BF16/FP16 tile, 64 KB for FP32) takes, we estimate, microseconds. We have not observed a failure on hardware; this is a conformance fix.

Steps/Code to reproduce bug

Found by checking a reduced extract of this sequence (a 4 KB linear bulk copy instead of the tensor-map load), compiled with CUDA 12.9 for sm_90a, against the PTX ISA. It has not been reproduced on hardware. A targeted probe would delay one non-leader warp before its first poll, use bounded waits, and compare against a version that joins first.

Expected behavior

All threads have finished waiting on tma_mbar before it is invalidated. The fix is to move the existing barrier above the invalidation at both sites:

ptx::mbarrier_wait_parity(&tma_mbar, 0);
__syncthreads();
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);

TE already does this elsewhere, for example in cast/mxfp8/dequantize_mxfp8.cuh (sm_100 path) and fused_attn/flash_attn.cu, which reach a __syncthreads() before invalidating their barriers. tma_mbar is not used after the invalidation, so no other change is needed.

Environment overview

Source analysis at 5e99464; no runtime environment involved.

Device details

Hopper (sm_90a) code path.

Additional context

Reach: the grouped FP8 block-scaling quantize kernels on Hopper (2D, 1D columnwise and both, and 1D rowwise with dbias through bgrad_group_quantize). te.ops.GroupedLinear reaches them by default on Hopper with cuBLASLt 13.6 or later and a block-scaling recipe; the te.pytorch.GroupedLinear module only with the opt-in use_grouped_tensor=True (or the deprecated NVTE_GROUPED_LINEAR_USE_FUSED_GROUPED_GEMM=1).

Dominant language
Python
Stars
3.6k
Forks
851
Avg merge
4d 14h
Merged PRs (30d)
54

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/TransformerEngine

All issues in NVIDIA/TransformerEngine

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.