Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[BUG] Fused MoE auxiliary loss accepts shapes that overflow int offsets

Open
#3,568 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 2 days

@arcusbuilds is already working on this.

Since Sep 24, 2026.

  • #3569 by @arcusbuilds — open

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
72/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
cpp

Research direction

Start in transformer_engine/common/fused_router/fused_moe_aux_loss.cu by tracing the index calculation and the forward, graph-safe forward, and backward launchers. Add validation before each launch for shapes whose flattened size exceeds the 32-bit offset limit, with an explanatory error; done means all three launchers reject the reported shape safely.

Written by the indexing model from the issue text.

Description

Describe the bug

The fused MoE auxiliary loss kernels calculate flattened positions as row * num_cols using a 32-bit int. The forward, graph-safe forward, and backward launchers accept shapes for which num_rows * num_cols exceeds INT_MAX. Those positions can overflow before the kernels use them to access memory.

Steps to reproduce

Inspect the index calculation and launchers in fused_moe_aux_loss.cu. For num_rows = 8_388_609 and num_cols = 256, the product is 2_147_483_904, which exceeds INT_MAX.

I found this through source inspection and have not run the oversized kernel, since an overflowing offset may access memory outside the tensor.

Expected behavior

Reject the shape before launching any of the three kernels, with an error explaining the 32-bit offset limit.

Environment

Source commit: 21a066d257afeaf8039325f58df9a62587517976. Runtime reproduction has not been run.

Dominant language
Python
Stars
3.6k
Forks
851
Avg merge
5d 1h
Merged PRs (30d)
52

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/TransformerEngine

All issues in NVIDIA/TransformerEngine

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.