Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on it

Đang mở Phù hợp với người mới
#3,647 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 2 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức phù hợp với người mới
82/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
python
Lĩnh vực
performance

Hướng nghiên cứu

Đọc transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh, tập trung vào group_block_scaled_2d_tma_kernel và group_block_scaled_1d_tma_kernel tại hai chuỗi barrier được trích dẫn. Di chuyển __syncthreads() hiện có lên trước thao tác vô hiệu hóa của leader ở cả hai vị trí, sau đó chạy các bài kiểm thử FP8 blockwise nhóm liên quan; hoàn tất khi tất cả các luồng đều đồng bộ trước khi một trong hai barrier bị vô hiệu hóa.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Describe the bug

In transformer_engine/common/cast/fp8_blockwise/group_quantize_fp8_blockwise.cuh (main 5e99464), both TMA kernels (group_block_scaled_2d_tma_kernel and group_block_scaled_1d_tma_kernel) end their TMA load like this (lines 343–345, and the same at 630–632):

ptx::mbarrier_wait_parity(&tma_mbar, 0);
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);
__syncthreads();

Every thread polls tma_mbar, and thread 0 invalidates it before the block joins. Nothing makes the other threads finish polling first. Warps execute independently, so another warp may not have started its poll yet. Also, mbarrier.try_wait may return after a system-dependent time limit and be retried (§9.7.15.16.19), so a lagging thread can poll again after the invalidation. Lanes 1–31 of warp 0 are not covered either: the wait loop is per thread, so lanes can diverge.

PTX ISA 9.4 (§9.7.15.16.13, mbarrier.inval): "Performing any mbarrier operation except mbarrier.init on a memory location that does not contain a valid mbarrier object, results in undefined behaviour." The ISA's own examples put a barrier before @t0 mbarrier.inval.

A late poll on the invalidated word is undefined. If invalidation changes the word, that thread may never see the phase complete, and the block hangs at the __syncthreads(). The window is narrow: the other warps start polling within cycles, while the load (32 KB for a BF16/FP16 tile, 64 KB for FP32) takes, we estimate, microseconds. We have not observed a failure on hardware; this is a conformance fix.

Steps/Code to reproduce bug

Found by checking a reduced extract of this sequence (a 4 KB linear bulk copy instead of the tensor-map load), compiled with CUDA 12.9 for sm_90a, against the PTX ISA. It has not been reproduced on hardware. A targeted probe would delay one non-leader warp before its first poll, use bounded waits, and compare against a version that joins first.

Expected behavior

All threads have finished waiting on tma_mbar before it is invalidated. The fix is to move the existing barrier above the invalidation at both sites:

ptx::mbarrier_wait_parity(&tma_mbar, 0);
__syncthreads();
if (leading_thread) ptx::mbarrier_invalid(&tma_mbar);

TE already does this elsewhere, for example in cast/mxfp8/dequantize_mxfp8.cuh (sm_100 path) and fused_attn/flash_attn.cu, which reach a __syncthreads() before invalidating their barriers. tma_mbar is not used after the invalidation, so no other change is needed.

Environment overview

Source analysis at 5e99464; no runtime environment involved.

Device details

Hopper (sm_90a) code path.

Additional context

Reach: the grouped FP8 block-scaling quantize kernels on Hopper (2D, 1D columnwise and both, and 1D rowwise with dbias through bgrad_group_quantize). te.ops.GroupedLinear reaches them by default on Hopper with cuBLASLt 13.6 or later and a block-scaling recipe; the te.pytorch.GroupedLinear module only with the opt-in use_grouped_tensor=True (or the deprecated NVTE_GROUPED_LINEAR_USE_FUSED_GROUPED_GEMM=1).

Ngôn ngữ chính
Python
Star
3.6k
Fork
851
Merge trung bình
4 ngày 20 giờ
Pull request đã merge (30 ngày)
56

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của NVIDIA/TransformerEngine

Tất cả issue của NVIDIA/TransformerEngine

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.