Multi-tensor swizzle kernels fail with "too many resources requested for launch" (missing __launch_bounds__)
Maintainer thường phản hồi trong vòng 2 ngày
Một pull request liên quan đã được merge.
- #3622 của @ravimajeti — đã merge
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 25/100
Hướng nghiên cứu
The four multi-tensor kernels live in transformer_engine/common/swizzle/swizzle.cu and are launched from launch_multi_tensor_swizzle_scaling_factors (the line in the report's stack trace); #2076 added launch_bounds(TB_DIM * TB_DIM) to the single-tensor kernels, so compare those declarations first. Done means the row, row and col variants stay within 64 registers at 1024 threads and the MultiTensorSwizzleTestSuite cases pass when run via tests/cpp/build/operator/test_operator --gtest_filter='MultiTensorSwizzleTestSuite'. Note PR #3622 is already open against this issue, so coordinate before starting.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the bug
The multi-tensor swizzle kernels in transformer_engine/common/swizzle/swizzle.cu (multi_tensor_swizzle_{row,col}_scaling_kernel and multi_tensor_unswizzle_{row,col}_scaling_kernel) are launched with 1024 threads per block (dim3 block_size(TB_DIM, TB_DIM)), but unlike every other swizzle kernel they have no __launch_bounds__(TB_DIM * TB_DIM). Without it, ptxas may give a thread more than 64 registers. 1024 threads × more than 64 registers exceeds the 64K registers available to a block, so the launch fails with:
CUDA Error: too many resources requested for launch
#2076 added __launch_bounds__ to the single-tensor swizzle kernels; the multi-tensor kernels added a week earlier in #2019 were not included.
Whether it fails depends on the GPU arch and CUDA version, because the register count does. Measured with cuobjdump --dump-resource-usage on swizzle.cu compiled with the TE build flags (registers per thread; > 64 cannot launch with 1024 threads):
| Kernel | CUDA 12.8 sm_90 | CUDA 12.8 sm_100 | CUDA 12.8 sm_120 | CUDA 13.4 sm_90 | CUDA 13.4 sm_100 | CUDA 13.4 sm_120 |
|---|---|---|---|---|---|---|
multi_tensor_swizzle_row_scaling_kernel<int4> |
89 | 89 | 96 | 61 | 50 | 56 |
multi_tensor_swizzle_row_scaling_kernel<int2> |
32 | 61 | 71 | 32 | 61 | 70 |
multi_tensor_swizzle_col_scaling_kernel<int4> |
89 | 99 | 99 | 64 | 56 | 61 |
| other multi-tensor (un)swizzle variants | ≤ 56 | ≤ 40 | ≤ 48 | ≤ 55 | ≤ 40 | ≤ 48 |
So on sm_120 it fails with both CUDA versions, and with CUDA 12.8 (the minimum for Blackwell) it should also fail on sm_90 and sm_100.
Steps/Code to reproduce bug
RTX 5090 (CC 12.0), TE built with NVTE_CUDA_ARCHS=120, CUDA 12.8:
tests/cpp/build/operator/test_operator --gtest_filter='*MultiTensorSwizzleTestSuite*'
3 cases fail, all routed to multi_tensor_swizzle_row_scaling_kernel<int4>:
[ FAILED ] OperatorTest/MultiTensorSwizzleTestSuite.TestMultiTensorSwizzle/n2_M128_K1024_row
[ FAILED ] OperatorTest/MultiTensorSwizzleTestSuite.TestMultiTensorSwizzle/n3_M256_K4096_row
[ FAILED ] OperatorTest/MultiTensorSwizzleTestSuite.TestMultiTensorSwizzle/n2_M128_K8192_row
C++ exception with description "transformer_engine/common/swizzle/swizzle.cu:1360 in function launch_multi_tensor_swizzle_scaling_factors: CUDA Error: too many resources requested for launch" thrown in the test body.
(n2_M128_K1024_row is meant to cover the narrow-K kernel, which needs 128 KB of shared memory. On a 99 KiB GPU the dispatcher correctly falls back to the regular multi-tensor kernel, which then fails to launch.)
The row<int2> and col<int4> variants are not reached by the current test shapes, e.g. {2, 128, 4352, true} (row, vec_load_size = 2) and {2, 512, 4096, false} (col, vec_load_size = 4) would cover them.
Expected behavior
The multi-tensor swizzle succeeds for all shapes, like the single-tensor path.
Proposed fix
Add __launch_bounds__(TB_DIM * TB_DIM) to the four multi-tensor kernels, matching the other swizzle kernels, and add test shapes that reach the row<int2> and col<int4> variants. I'm happy to open a PR.
Environment overview
- Environment location: Docker on vast.ai
- Method of Transformer Engine install: from source (
mainat 5759fa0f) - Docker image: Ubuntu 24.04 with
/venv/main
Environment details
- OS version: Ubuntu 24.04
- PyTorch version: 2.11.0+cu128
- Python version: 3.12
- Transformer Engine version: 2.21.0.dev0+5759fa0f
- CUDA version: 12.8 (register table also with 13.4)
- CUDNN version: 9.19
Device details
- GPU model: NVIDIA GeForce RTX 5090 (CC 12.0), driver 580.95.05
- Ngôn ngữ chính
- Python
- Star
- 3.6k
- Fork
- 844
- Merge trung bình
- 4 ngày 15 giờ
- Pull request đã merge (30 ngày)
- 51
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/TransformerEngine
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA/TransformerEngine#3636 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itCó thể đã có người làm @yuweih205 đã nhận 30 ngày trước. Đang mởattention
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
NVIDIA/TransformerEngine#3481 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Increase MAX_TENSOR_NUMĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
NVIDIA/TransformerEngine#2189 · 7 bình luận · 5 reaction ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 50/100
NVIDIA/TransformerEngine#3645 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
enhancement
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
NVIDIA/TransformerEngine#3644 ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của NVIDIA/TransformerEngine
Issue tương tự
-
area/install reliability
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 83/100
FluidNumerics/fluid-walk-blocker#191 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
TransformerLensOrg/TransformerLens#1868 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 85/100
climate-analytics-lab/jax-gcm#1057 ·
Maintainer thường phản hồi trong vòng 1 ngày