[PyTorch] Request a migration path for downstream `_GroupedLinear` integrations after the 2.17 signature change
Maintainer thường phản hồi trong vòng 2 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Lĩnh vực
- backend, machine-learning
Hướng nghiên cứu
Bắt đầu với transformer_engine/pytorch/module/grouped_linear.py và kiểm tra chữ ký của _GroupedLinear.forward cũng như cơ chế dispatch của linear_fn quanh commit f8bda5d0. Chạy pytest target Megatron-Bridge được liệt kê để tái hiện lỗi len(False), sau đó xem xét các ví dụ tương thích ModelOpt và Megatron-Bridge được liên kết. Hoàn tất có nghĩa là hành vi migration hoặc transition được hỗ trợ đã được ghi lại và được bao phủ bởi một regression test cho các lệnh gọi wrapped linear_fn.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the bug
Transformer Engine 2.17 intentionally changed the private PyTorch _GroupedLinear autograd call layout while making GroupedLinear graph-safe.
TE 2.16:
forward(ctx, inp, non_tensor_args, *weights_and_biases)
# non_tensor_args[0] is m_splits
TE 2.17:
forward(ctx, inp, m_splits, non_tensor_args, *weights_and_biases)
# m_splits is now a separate int64 tensor
# non_tensor_args[0] is use_bias
The change was introduced by f8bda5d0. The public GroupedLinear.forward accepts both list and tensor split inputs, while the internal linear_fn dispatch boundary changed positionally.
We understand that _GroupedLinear is private and that downstream users cannot assume its signature will remain stable. In practice, however, there does not appear to be a public functional or interception API for these use cases, so some downstream integrations currently use this boundary to intercept grouped GEMMs or invoke them with externally owned weights. Upgrading TE alone therefore causes deterministic runtime failures in those integrations.
Two downstream examples are affected:
-
NVIDIA ModelOpt intercepts the grouped-linear function to quantize its input and weights. Its compatibility logic finds
non_tensor_argsin the signature and readsnon_tensor_args[0]asm_splits. With TE 2.17, that value isuse_bias, so calibration fails with:File ".../transformer_engine/pytorch/module/grouped_linear.py", line 1788, in forward out, new_workspaces = linear_fn( File ".../modelopt/torch/quantization/plugins/transformer_engine.py", line 178, in te_grouped_quantized_linear_fn num_gemms = len(args[sig_params.index("non_tensor_args") - ctx_offset][0]) TypeError: object of type 'bool' has no len()Tracking issue: NVIDIA/Model-Optimizer#1940
-
Megatron-Bridge directly calls
_GroupedLinear.apply/forwardwith externally owned grouped adapter weights. The TE 2.16 layout passesx, non_tensor_args, weights..., biases.... Under TE 2.17, the same call is shifted: the old tuple is interpreted asm_splits, the first weight is interpreted asnon_tensor_args, and the remaining weights and biases are mispartitioned. The downstream compatibility fix is NVIDIA-NeMo/Megatron-Bridge#4721.
Steps/Code to reproduce bug
The failure is reproduced by changing only the TE pin in Megatron-Bridge PR #4696:
- Passing baseline:
2.16.0+d64bc14datd64bc14dc87eb658ab98839e4b7687595ee53e2d - Failing version:
2.17.0+2e559f06at2e559f062497bef768dfbe9d7e45548fadeca80a - ModelOpt remains pinned at
nvidia-modelopt==0.44.0rc5
Run:
uv run python -m pytest -s -x \
tests/functional_tests/test_groups/quantization/models/qwen/test_qwen3_moe_quantization_workflow.py::TestQwen3MoeQuantizationWorkflow::test_qwen3_moe_quantization_and_generation_with_expert_parallelism
The first calibration forward fails at the len(False) exception shown above. The same failure occurs on both GB200 and H100 runners. Full failing job: Megatron-Bridge GitHub Actions.
The signature transition can also be confirmed directly:
import inspect
from transformer_engine.pytorch.module.grouped_linear import _GroupedLinear
print(inspect.signature(_GroupedLinear.forward))
Requested guidance / compatibility support
We will update the affected downstream integrations to handle the new layout. To make that migration less disruptive, would the TE maintainers consider publishing a TE 2.17.x patch release with a short-lived transition path for both grouped-linear call layouts?
# Legacy layout used through TE 2.16
(ctx, inp, non_tensor_args, *weights_and_biases)
# Graph-safe layout introduced in TE 2.17
(ctx, inp, m_splits, non_tensor_args, *weights_and_biases)
If practical, the compatibility path could cover both grad-enabled _GroupedLinear.apply and direct no-grad _GroupedLinear.forward dispatch, preserve the correct backward arity, and emit a deprecation warning for the legacy form before it is removed in a later feature release. We are not asking for indefinite compatibility for the private API.
Because wrappers such as ModelOpt intercept the linear_fn boundary before _GroupedLinear.forward executes, compatibility may need to be provided at that dispatch/hook boundary rather than only inside the autograd function. A regression test using a wrapped linear_fn would cover this downstream use case.
If dual-layout support at this private boundary is not feasible, guidance on the intended supported approach—or an equivalent stable functional/interception API—would be equally helpful. Public GroupedLinear owns its parameters, so it is not currently a drop-in replacement for callers that need to supply externally owned grouped weights.
Environment overview (please complete the following information)
- Environment location: GitHub Actions, Docker, GCP GPU runners
- Transformer Engine install: Git source pin resolved by
uv - Passing TE:
2.16.0+d64bc14d - Failing TE:
2.17.0+2e559f06 - Downstream packages: Megatron-Bridge PR #4696 and
nvidia-modelopt==0.44.0rc5
Environment details
- OS: Linux container
- Python: 3.12.3
- Transformer Engine: versions and exact commits listed above
- The full container and runner information is recorded in the linked GitHub Actions job.
Device details
- Reproduced on GB200 and H100 CI runners.
Additional context
The old and new contracts are straightforward for downstream code to distinguish when it owns the call: prefer the explicit m_splits parameter when present, otherwise use non_tensor_args[0]. The harder compatibility case is an interceptor that TE calls through the changed positional boundary, which is why a short TE-side transition window would be valuable.
This report concerns only the grouped-linear signature transition. It is independent of other TE 2.17 packaging or import behavior changes.
- Ngôn ngữ chính
- Python
- Star
- 3.6k
- Fork
- 851
- Merge trung bình
- 4 ngày 20 giờ
- Pull request đã merge (30 ngày)
- 56
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/TransformerEngine
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on itĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA/TransformerEngine#3647 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarCó thể đã có người làm @sanjana658 đã nhận 1 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA/TransformerEngine#3636 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itCó thể đã có người làm @yuweih205 đã nhận 32 ngày trước. Đang mởattention
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
NVIDIA/TransformerEngine#3481 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Increase MAX_TENSOR_NUMĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
NVIDIA/TransformerEngine#2189 · 7 bình luận · 5 reaction ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Fused gemm + comm for CP A2A on BlackwellCó thể đã có người làm @cyanguwa đã nhận hôm nay. Đang mở2.22 attention
NVIDIA/TransformerEngine#3664 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của NVIDIA/TransformerEngine
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
mishraprafful/multihull#150 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 66/100
python-caldav/caldav#735 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
mealie-recipes/mealie#8682 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
good first issue lane:repo
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100