[BUG] Fused MoE auxiliary loss accepts shapes that overflow int offsets
Maintainers usually reply within 2 days
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 72/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- cpp
- Domain
- machine-learning
Research direction
Start in transformer_engine/common/fused_router/fused_moe_aux_loss.cu by tracing the index calculation and the forward, graph-safe forward, and backward launchers. Add validation before each launch for shapes whose flattened size exceeds the 32-bit offset limit, with an explanatory error; done means all three launchers reject the reported shape safely.
Written by the indexing model from the issue text.
Description
Describe the bug
The fused MoE auxiliary loss kernels calculate flattened positions as row * num_cols using a 32-bit int. The forward, graph-safe forward, and backward launchers accept shapes for which num_rows * num_cols exceeds INT_MAX. Those positions can overflow before the kernels use them to access memory.
Steps to reproduce
Inspect the index calculation and launchers in fused_moe_aux_loss.cu. For num_rows = 8_388_609 and num_cols = 256, the product is 2_147_483_904, which exceeds INT_MAX.
I found this through source inspection and have not run the oversized kernel, since an overflowing offset may access memory outside the tensor.
Expected behavior
Reject the shape before launching any of the three kernels, with an error explaining the 32-bit offset limit.
Environment
Source commit: 21a066d257afeaf8039325f58df9a62587517976. Runtime reproduction has not been run.
- Dominant language
- Python
- Stars
- 3.6k
- Forks
- 851
- Avg merge
- 5d 1h
- Merged PRs (30d)
- 52
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/TransformerEngine
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on itOpen
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA/TransformerEngine#3647 ·
Maintainers usually reply within 2 days
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarPossibly taken @sanjana658 claimed this 1 day ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA/TransformerEngine#3636 · 2 comments ·
Maintainers usually reply within 2 days
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itPossibly taken @yuweih205 claimed this 32 days ago. Openattention
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
NVIDIA/TransformerEngine#3481 · 4 comments ·
Maintainers usually reply within 2 days
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
NVIDIA/TransformerEngine#2189 · 7 comments · 5 reactions ·
Maintainers usually reply within 2 days
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration buildsPossibly taken @ksivaman claimed this today. Open
Difficulty 4/5 3-5 days Newbie friendliness 50/100
NVIDIA/TransformerEngine#3645 · 1 comment · 1 assignee ·
Maintainers usually reply within 2 days
All issues in NVIDIA/TransformerEngine
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
NVIDIA/earth2studio#1241 ·
Maintainers usually reply within 3 days
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0Opendocumentation priority: low size: S
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
Maintainers usually reply within 1 day
-
bug help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Maintainers usually reply within 1 day
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 65/100
ansys/pydpf-core#3547 ·
Maintainers usually reply within 1 day
-
good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
OktoLabsAI/okto-pulse#114 ·
Maintainers usually reply within 1 day