Support exporting MXFP8 values and scales from TE quantized tensors
Maintainer thường phản hồi trong vòng 2 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Lĩnh vực
- machine-learning
Hướng nghiên cứu
Start by reading the TE quantized-tensor APIs and storage handling for MXFP8; the issue does not name specific files, tests, or entry points. Define and implement a supported export interface for values, scales, and interpretation metadata, covering both existing quantized tensors and newly quantized tensors. Done means callers can use both workflows without relying on internal storage conventions, with copying, layout conversion, and storage lifetime behavior documented and tested.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Is your feature request related to a problem? Please describe.
RL frameworks frequently synchronize training weights with inference engines. For low-precision inference, this requires exporting quantized values and their associated scales.
There are two common workflows:
- Reuse existing quantized weights. When training already stores compatible, MXFP8 params, export their values and scales directly.
- Quantize weights for inference. When training stores BF16 parameters or inference needs different quantization, quantize the exported weights, then extract the resulting values and scales.
Both workflows need access to TE’s quantized tensor components. Today, downstream frameworks handle storage details such as padding, scale layouts, and byte interpretation themselves.
For example, https://github.com/NVIDIA-NeMo/RL/pull/3908 extracts native MXFP8 storage through TE metadata, while Miles accesses internal buffers after TE quantization.
Describe the solution you'd like
Provide a supported way to export a TE quantized tensor’s values, scales, and the metadata needed to interpret them outside TE, initially for MXFP8.
This should support both existing quantized training parameters and newly quantized tensors. The goal is to let downstream integrations consume these components without depending on TE’s internal storage conventions.
Where compatible quantized storage already exists, export should preserve that representation without unnecessary dequantization and requantization. Copying, layout conversion, and storage lifetime behavior should be clear to callers.
Model-level conversion and distributed mappings would remain in tools such as Megatron Bridge. Synchronization, transport, and inference-specific loading would remain in downstream RL frameworks.
Describe alternatives you've considered
- Read TE metadata or internal buffers downstream. This works today but requires each integration to understand and maintain TE-specific extraction logic.
- Always convert to BF16 and requantize. This adds unnecessary work when compatible quantized storage already exists.
Additional context
Examples of relevant work:
- NeMo RL #3908: exports existing native MXFP8 training weights.
- Miles MXFP8 helper: extracts components after quantizing weights with TE.
- Megatron-Bridge #5917: model-level native MXFP8 export support.
- Verl #8129: proposed TE-based MXFP8 quantization during rollout synchronization.
- Slime low-precision documentation: describes BF16 conversion and FP8 quantization during synchronization.
- Ngôn ngữ chính
- Python
- Star
- 3.6k
- Fork
- 851
- Merge trung bình
- 4 ngày 15 giờ
- Pull request đã merge (30 ngày)
- 51
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/TransformerEngine
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA/TransformerEngine#3636 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itCó thể đã có người làm @yuweih205 đã nhận 31 ngày trước. Đang mởattention
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
NVIDIA/TransformerEngine#3481 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Increase MAX_TENSOR_NUMĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
NVIDIA/TransformerEngine#2189 · 7 bình luận · 5 reaction ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 50/100
NVIDIA/TransformerEngine#3645 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 54/100
NVIDIA/TransformerEngine#3640 · 8 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của NVIDIA/TransformerEngine
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
RedHatQE/mtv-api-tests#721 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 85/100
pytest-dev/pluggy#757 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 85/100
NousResearch/hermes-agent#134960 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
HTML backend: `<br>` leaks the internal sentinel U+E000 into list items, headings and captionsCó thể đã có người làm @morten-lagabote đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 67/100
docling-project/docling#4671 ·
Maintainer thường phản hồi trong vòng 1 ngày