Hybrid quantization follow-ups
メンテナーはふだん 2 日以内に返信
@negvet がすでに取り組んでいます。
2026年6月30日 から。
評価
この issue はまだ評価されていません。
説明
Tracking follow-ups from PR #2817. Detailed implementation notes are already in code comments / xfails.
Recipe / validation
- Validate per-GEMM scaling-mode compatibility for hybrid qfactories.
- See
HybridQuantizer._get_compatible_recipe()intransformer_engine/pytorch/tensor/hybrid_tensor.py.
- See
- Support delayed-scaling requests inside
HybridQuantizer.- See
HybridQuantizer.__init__()in `transformer_engine/pytorch/tensor/hybrid_tensor.py
- See
- Support producing both directional representations at the fused normalization boundary (transformer_engine/pytorch/ops/basic/layer_norm.py)
Performance / kernels
- Fused kernels: one read -> two writes.
- Introduce a policy for when to materialize a tensor for
columnwise_source="rowwise_dequantized". - Support module-/role-specific CustomRecipe quantization alignment. The current recipe-global quantization_alignment must use the maximum requirement across all factory outputs, which can overpad lower-alignment MXFP8/FP8 layers in mixed-format models. Derive alignment from canonical cached module quantizers and propagate it to the corresponding Fp8Padding without speculative qfactory calls.
- Support HybridQuantizer with standard Userbuffers.
- Add native columnwise-only per-tensor FP8 quantization for Hopper, covering both CurrentScaling and DelayedScaling. The legacy per-tensor cast-transpose kernel currently requires both rowwise and columnwise output buffers.
TP/SP
- Add native
HybridQuantizerdispatch togather_along_first_dim. - Preserve existing SP amax-reduction semantics and cover them in distributed tests.
FSDP2
- Optimize hybrid FSDP2 communication buffers.
- See
HybridQuantizedTensor.fsdp_pre_all_gather()intransformer_engine/pytorch/tensor/hybrid_tensor.py.
- See
- Fix
HybridFloat8BlockScalingFSDP2 xfail.- See
_HYBRID_FLOAT8_BLOCK_FSDP2_XFAIL_REASONintests/pytorch/distributed/fsdp2_tests/conftest.py.
- See
- Add NVFP4 hybrid sub-storage FSDP2 hooks.
- See
TestHybridFsdpPreAllGatherProtocol.test_nvfp4_sub_storage_raises_on_pre_all_gather()intests/pytorch/test_hybrid_quantization.py. - Non-hybrid FSDP alignment: reuse fsdp extract buffers and fsdp_assign_gather for the current base tensor class (non-hybrid)
- See
- Support Hybrid sub-storages that fall back to a high-precision tensor for an unquantizable local shard (TransformerEngine/transformer_engine/pytorch/tensor/hybrid_tensor.py).
GEMM / quantization
- Support
HybridQuantizer/IdentityQuantizeras GEMM output quantizers.- See
_reject_unsupported_output_quantizer()intransformer_engine/pytorch/cpp_extensions/gemm.py.
- See
GroupedLinear / grouped storage
- Support
IdentityQuantizerandHybridQuantizerwithGroupedLinear(single_grouped_weight=True).
Distributed optimizer / Megatron
- Support per-block hybrid sub-quantizers in
quantize_master_weights. - Hybrid distributed optimizer: partial-master support with columnwise_source="original". The current one-payload distributed-optimizer path cannot preserve both independently quantized Hybrid directions when each rank owns only a partial master weight. A safety guard rejects this configuration unless full-master data or an FSDP Hybrid shard containing both directions is provided. Unblock by adding a two-payload sharding/all-gather contract for Hybrid rowwise and columnwise storage, including their scale/amax metadata, then relax _validate_hybrid_partial_master_policy.
- Hybrid distributed optimizer with columnwise_source="rowwise_dequantized": reconstruct the column only after the rowwise update/all-gather.
- Complete Megatron-LM
quantized_model_init + --fp{4,8}-param-gather + dist opt.- See PR #2817 integration notes.
- Complete Megatron-FSDP +
--fp{4,8}-param-gather.- See PR #2817 integration notes.
- Complete Torch FSDP2 +
--fp{4,8}-param-gather.- See PR #2817 integration notes and
tests/pytorch/distributed/fsdp2_tests/.
- See PR #2817 integration notes and
Activation recompute
- Investigate vanilla
torch.utils.checkpoint(use_reentrant=False)with TE weight-workspace cache.- See xfails in
TestHybridActivationRecomputeintests/pytorch/test_hybrid_quantization.py. te.checkpointpath is already covered and works.
- See xfails in
Validation
- Convergence validation of base non-hybrid recipes.
- 主要言語
- Python
- スター
- 3.6k
- フォーク
- 851
- 平均マージ
- 5日 1時間
- マージ済み PR(30日)
- 52
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/TransformerEngine のほかの issue
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on itオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3647 ·
メンテナーはふだん 2 日以内に返信
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar対応中かも @sanjana658 が 1 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3636 · コメント 2 件 ·
メンテナーはふだん 2 日以内に返信
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it対応中かも @yuweih205 が 32 日前に担当しました。 オープンattention
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
NVIDIA/TransformerEngine#3481 · コメント 4 件 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
NVIDIA/TransformerEngine#2189 · コメント 7 件 · リアクション 5 件 ·
メンテナーはふだん 2 日以内に返信
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration builds対応中かも @ksivaman が今日担当しました。 オープン
難易度 4/5 3〜5日 初心者へのやさしさ 50/100
NVIDIA/TransformerEngine#3645 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 2 日以内に返信
NVIDIA/TransformerEngine の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
NVIDIA/earth2studio#1241 ·
メンテナーはふだん 3 日以内に返信
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0オープンdocumentation priority: low size: S
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
メンテナーはふだん 1 日以内に返信
-
bug help wanted
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
メンテナーはふだん 1 日以内に返信
-
documentation
難易度 1/5 1時間未満 初心者へのやさしさ 65/100
ansys/pydpf-core#3547 ·
メンテナーはふだん 1 日以内に返信
-
good first issue
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
OktoLabsAI/okto-pulse#114 ·
メンテナーはふだん 1 日以内に返信