Support exporting MXFP8 values and scales from TE quantized tensors
维护者通常 2 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 35/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
调研方向
Start by reading the TE quantized-tensor APIs and storage handling for MXFP8; the issue does not name specific files, tests, or entry points. Define and implement a supported export interface for values, scales, and interpretation metadata, covering both existing quantized tensors and newly quantized tensors. Done means callers can use both workflows without relying on internal storage conventions, with copying, layout conversion, and storage lifetime behavior documented and tested.
由索引模型根据 Issue 内容生成。
描述
Is your feature request related to a problem? Please describe.
RL frameworks frequently synchronize training weights with inference engines. For low-precision inference, this requires exporting quantized values and their associated scales.
There are two common workflows:
- Reuse existing quantized weights. When training already stores compatible, MXFP8 params, export their values and scales directly.
- Quantize weights for inference. When training stores BF16 parameters or inference needs different quantization, quantize the exported weights, then extract the resulting values and scales.
Both workflows need access to TE’s quantized tensor components. Today, downstream frameworks handle storage details such as padding, scale layouts, and byte interpretation themselves.
For example, https://github.com/NVIDIA-NeMo/RL/pull/3908 extracts native MXFP8 storage through TE metadata, while Miles accesses internal buffers after TE quantization.
Describe the solution you'd like
Provide a supported way to export a TE quantized tensor’s values, scales, and the metadata needed to interpret them outside TE, initially for MXFP8.
This should support both existing quantized training parameters and newly quantized tensors. The goal is to let downstream integrations consume these components without depending on TE’s internal storage conventions.
Where compatible quantized storage already exists, export should preserve that representation without unnecessary dequantization and requantization. Copying, layout conversion, and storage lifetime behavior should be clear to callers.
Model-level conversion and distributed mappings would remain in tools such as Megatron Bridge. Synchronization, transport, and inference-specific loading would remain in downstream RL frameworks.
Describe alternatives you've considered
- Read TE metadata or internal buffers downstream. This works today but requires each integration to understand and maintain TE-specific extraction logic.
- Always convert to BF16 and requantize. This adds unnecessary work when compatible quantized storage already exists.
Additional context
Examples of relevant work:
- NeMo RL #3908: exports existing native MXFP8 training weights.
- Miles MXFP8 helper: extracts components after quantizing weights with TE.
- Megatron-Bridge #5917: model-level native MXFP8 export support.
- Verl #8129: proposed TE-based MXFP8 quantization during rollout synchronization.
- Slime low-precision documentation: describes BF16 conversion and FP8 quantization during synchronization.
- 主要语言
- Python
- 星标
- 3.6k
- 派生
- 851
- 平均合并
- 4 天 15 小时
- 30 天内合并 PR
- 51
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/TransformerEngine 的其他 Issue
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar未关闭
难度 2/5 1-3 小时 新手友好度 82/100
NVIDIA/TransformerEngine#3636 ·
维护者通常 2 天内回复
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it可能已有人在做 @yuweih205 于 31 天前认领。 未关闭attention
难度 2/5 1-3 小时 新手友好度 85/100
NVIDIA/TransformerEngine#3481 · 4 条评论 ·
维护者通常 2 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 68/100
NVIDIA/TransformerEngine#2189 · 7 条评论 · 5 个 reaction ·
维护者通常 2 天内回复
-
难度 4/5 3-5 天 新手友好度 50/100
NVIDIA/TransformerEngine#3645 ·
维护者通常 2 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 54/100
NVIDIA/TransformerEngine#3640 · 8 条评论 ·
维护者通常 2 天内回复
查看 NVIDIA/TransformerEngine 的全部 Issue
相似的 Issue
-
难度 1/5 1-3 小时 新手友好度 85/100
pytest-dev/pluggy#757 ·
维护者通常 1 天内回复
-
难度 1/5 1-3 小时 新手友好度 85/100
NousResearch/hermes-agent#134960 ·
维护者通常 1 天内回复
-
HTML backend: `<br>` leaks the internal sentinel U+E000 into list items, headings and captions可能已有人在做 @morten-lagabote 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 67/100
docling-project/docling#4671 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 70/100
维护者通常 1 天内回复
-
good first issue hacktoberfest infra
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复