Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on it

Aperta Adatta ai principianti
#3,647 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 2 giorni

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
82/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
performance

Direzione di ricerca

Leggi transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh, concentrandoti su group_block_scaled_2d_tma_kernel e group_block_scaled_1d_tma_kernel nelle due sequenze di barriere citate. Sposta il __syncthreads() esistente prima dell’invalidazione da parte del leader in entrambi i punti, quindi esegui i test pertinenti per FP8 blockwise raggruppato; il lavoro è completato quando tutti i thread si sincronizzano prima che venga invalidata una delle due barriere.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Describe the bug

In transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh (main 5e99464), both TMA kernels (group_block_scaled_2d_tma_kernel and group_block_scaled_1d_tma_kernel) end their TMA load like this (lines 343–345, and the same at 630–632):

ptx::mbarrier_wait_parity(&tma_mbar, 0);
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);
__syncthreads();

Every thread polls tma_mbar, and thread 0 invalidates it before the block joins. Nothing makes the other threads finish polling first. Warps execute independently, so another warp may not have started its poll yet. Also, mbarrier.try_wait may return after a system-dependent time limit and be retried (§9.7.15.16.19), so a lagging thread can poll again after the invalidation. Lanes 1–31 of warp 0 are not covered either: the wait loop is per thread, so lanes can diverge.

PTX ISA 9.4 (§9.7.15.16.13, mbarrier.inval): "Performing any mbarrier operation except mbarrier.init on a memory location that does not contain a valid mbarrier object, results in undefined behaviour." The ISA's own examples put a barrier before @t0 mbarrier.inval.

A late poll on the invalidated word is undefined. If invalidation changes the word, that thread may never see the phase complete, and the block hangs at the __syncthreads(). The window is narrow: the other warps start polling within cycles, while the load (32 KB for a BF16/FP16 tile, 64 KB for FP32) takes, we estimate, microseconds. We have not observed a failure on hardware; this is a conformance fix.

Steps/Code to reproduce bug

Found by checking a reduced extract of this sequence (a 4 KB linear bulk copy instead of the tensor-map load), compiled with CUDA 12.9 for sm_90a, against the PTX ISA. It has not been reproduced on hardware. A targeted probe would delay one non-leader warp before its first poll, use bounded waits, and compare against a version that joins first.

Expected behavior

All threads have finished waiting on tma_mbar before it is invalidated. The fix is to move the existing barrier above the invalidation at both sites:

ptx::mbarrier_wait_parity(&tma_mbar, 0);
__syncthreads();
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);

TE already does this elsewhere, for example in cast/mxfp8/dequantize_mxfp8.cuh (sm_100 path) and fused_attn/flash_attn.cu, which reach a __syncthreads() before invalidating their barriers. tma_mbar is not used after the invalidation, so no other change is needed.

Environment overview

Source analysis at 5e99464; no runtime environment involved.

Device details

Hopper (sm_90a) code path.

Additional context

Reach: the grouped FP8 block-scaling quantize kernels on Hopper (2D, 1D columnwise and both, and 1D rowwise with dbias through bgrad_group_quantize). te.ops.GroupedLinear reaches them by default on Hopper with cuBLASLt 13.6 or later and a block-scaling recipe; the te.pytorch.GroupedLinear module only with the opt-in use_grouped_tensor=True (or the deprecated NVTE_GROUPED_LINEAR_USE_FUSED_GROUPED_GEMM=1).

Lingua principale
Python
Stelle
3.6k
Fork
851
Merge medio
5g 1h
PR unite (30g)
52

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di NVIDIA/TransformerEngine

Tutte le issue di NVIDIA/TransformerEngine

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.