Support exporting MXFP8 values and scales from TE quantized tensors
メンテナーはふだん 2 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
調査の方向性
Start by reading the TE quantized-tensor APIs and storage handling for MXFP8; the issue does not name specific files, tests, or entry points. Define and implement a supported export interface for values, scales, and interpretation metadata, covering both existing quantized tensors and newly quantized tensors. Done means callers can use both workflows without relying on internal storage conventions, with copying, layout conversion, and storage lifetime behavior documented and tested.
索引モデルが issue の本文から書いたものです。
説明
Is your feature request related to a problem? Please describe.
RL frameworks frequently synchronize training weights with inference engines. For low-precision inference, this requires exporting quantized values and their associated scales.
There are two common workflows:
- Reuse existing quantized weights. When training already stores compatible, MXFP8 params, export their values and scales directly.
- Quantize weights for inference. When training stores BF16 parameters or inference needs different quantization, quantize the exported weights, then extract the resulting values and scales.
Both workflows need access to TE’s quantized tensor components. Today, downstream frameworks handle storage details such as padding, scale layouts, and byte interpretation themselves.
For example, https://github.com/NVIDIA-NeMo/RL/pull/3908 extracts native MXFP8 storage through TE metadata, while Miles accesses internal buffers after TE quantization.
Describe the solution you'd like
Provide a supported way to export a TE quantized tensor’s values, scales, and the metadata needed to interpret them outside TE, initially for MXFP8.
This should support both existing quantized training parameters and newly quantized tensors. The goal is to let downstream integrations consume these components without depending on TE’s internal storage conventions.
Where compatible quantized storage already exists, export should preserve that representation without unnecessary dequantization and requantization. Copying, layout conversion, and storage lifetime behavior should be clear to callers.
Model-level conversion and distributed mappings would remain in tools such as Megatron Bridge. Synchronization, transport, and inference-specific loading would remain in downstream RL frameworks.
Describe alternatives you've considered
- Read TE metadata or internal buffers downstream. This works today but requires each integration to understand and maintain TE-specific extraction logic.
- Always convert to BF16 and requantize. This adds unnecessary work when compatible quantized storage already exists.
Additional context
Examples of relevant work:
- NeMo RL #3908: exports existing native MXFP8 training weights.
- Miles MXFP8 helper: extracts components after quantizing weights with TE.
- Megatron-Bridge #5917: model-level native MXFP8 export support.
- Verl #8129: proposed TE-based MXFP8 quantization during rollout synchronization.
- Slime low-precision documentation: describes BF16 conversion and FP8 quantization during synchronization.
- 主要言語
- Python
- スター
- 3.6k
- フォーク
- 851
- 平均マージ
- 4日 15時間
- マージ済み PR(30日)
- 51
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/TransformerEngine のほかの issue
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
NVIDIA/TransformerEngine#3636 ·
メンテナーはふだん 2 日以内に返信
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it対応中かも @yuweih205 が 31 日前に担当しました。 オープンattention
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
NVIDIA/TransformerEngine#3481 · コメント 4 件 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
NVIDIA/TransformerEngine#2189 · コメント 7 件 · リアクション 5 件 ·
メンテナーはふだん 2 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 50/100
NVIDIA/TransformerEngine#3645 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 54/100
NVIDIA/TransformerEngine#3640 · コメント 8 件 ·
メンテナーはふだん 2 日以内に返信
NVIDIA/TransformerEngine の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
メンテナーはふだん 1 日以内に返信
-
SR_SECURITY_DESCRIPTOR.fromString drops the SACL when no DACL is present対応中かも @paul7436 が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
メンテナーはふだん 2 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
equinor/fmu-sumo-uploader#302 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
modelscope/evalscope#1821 ·
メンテナーはふだん 1 日以内に返信
-
Sanity on ansible-core devel fails: ignore-2.23.txt references the removed import-3.9 test対応中かも @yurnov が今日担当しました。 オープンneeds_triage
難易度 1/5 1時間未満 初心者へのやさしさ 91/100
ansible-collections/kubernetes.core#1275 ·
メンテナーはふだん 1 日以内に返信