Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Support exporting MXFP8 values and scales from TE quantized tensors

オープン
#3,644 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 2 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
活発
技術スタック
python, pytorch

調査の方向性

Start by reading the TE quantized-tensor APIs and storage handling for MXFP8; the issue does not name specific files, tests, or entry points. Define and implement a supported export interface for values, scales, and interpretation metadata, covering both existing quantized tensors and newly quantized tensors. Done means callers can use both workflows without relying on internal storage conventions, with copying, layout conversion, and storage lifetime behavior documented and tested.

索引モデルが issue の本文から書いたものです。

説明

enhancement

Is your feature request related to a problem? Please describe.

RL frameworks frequently synchronize training weights with inference engines. For low-precision inference, this requires exporting quantized values and their associated scales.

There are two common workflows:

  1. Reuse existing quantized weights. When training already stores compatible, MXFP8 params, export their values and scales directly.
  2. Quantize weights for inference. When training stores BF16 parameters or inference needs different quantization, quantize the exported weights, then extract the resulting values and scales.

Both workflows need access to TE’s quantized tensor components. Today, downstream frameworks handle storage details such as padding, scale layouts, and byte interpretation themselves.

For example, https://github.com/NVIDIA-NeMo/RL/pull/3908 extracts native MXFP8 storage through TE metadata, while Miles accesses internal buffers after TE quantization.

Describe the solution you'd like

Provide a supported way to export a TE quantized tensor’s values, scales, and the metadata needed to interpret them outside TE, initially for MXFP8.

This should support both existing quantized training parameters and newly quantized tensors. The goal is to let downstream integrations consume these components without depending on TE’s internal storage conventions.

Where compatible quantized storage already exists, export should preserve that representation without unnecessary dequantization and requantization. Copying, layout conversion, and storage lifetime behavior should be clear to callers.

Model-level conversion and distributed mappings would remain in tools such as Megatron Bridge. Synchronization, transport, and inference-specific loading would remain in downstream RL frameworks.

Describe alternatives you've considered

  • Read TE metadata or internal buffers downstream. This works today but requires each integration to understand and maintain TE-specific extraction logic.
  • Always convert to BF16 and requantize. This adds unnecessary work when compatible quantized storage already exists.

Additional context

Examples of relevant work:

主要言語
Python
スター
3.6k
フォーク
851
平均マージ
4日 15時間
マージ済み PR(30日)
51

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/TransformerEngine のほかの issue

NVIDIA/TransformerEngine の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。