Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Support exporting MXFP8 values and scales from TE quantized tensors

Đang mở
#3,644 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 2 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
35/100
Loại issue
Tính năng
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
python, pytorch
Lĩnh vực
machine-learning

Hướng nghiên cứu

Start by reading the TE quantized-tensor APIs and storage handling for MXFP8; the issue does not name specific files, tests, or entry points. Define and implement a supported export interface for values, scales, and interpretation metadata, covering both existing quantized tensors and newly quantized tensors. Done means callers can use both workflows without relying on internal storage conventions, with copying, layout conversion, and storage lifetime behavior documented and tested.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

enhancement

Is your feature request related to a problem? Please describe.

RL frameworks frequently synchronize training weights with inference engines. For low-precision inference, this requires exporting quantized values and their associated scales.

There are two common workflows:

  1. Reuse existing quantized weights. When training already stores compatible, MXFP8 params, export their values and scales directly.
  2. Quantize weights for inference. When training stores BF16 parameters or inference needs different quantization, quantize the exported weights, then extract the resulting values and scales.

Both workflows need access to TE’s quantized tensor components. Today, downstream frameworks handle storage details such as padding, scale layouts, and byte interpretation themselves.

For example, https://github.com/NVIDIA-NeMo/RL/pull/3908 extracts native MXFP8 storage through TE metadata, while Miles accesses internal buffers after TE quantization.

Describe the solution you'd like

Provide a supported way to export a TE quantized tensor’s values, scales, and the metadata needed to interpret them outside TE, initially for MXFP8.

This should support both existing quantized training parameters and newly quantized tensors. The goal is to let downstream integrations consume these components without depending on TE’s internal storage conventions.

Where compatible quantized storage already exists, export should preserve that representation without unnecessary dequantization and requantization. Copying, layout conversion, and storage lifetime behavior should be clear to callers.

Model-level conversion and distributed mappings would remain in tools such as Megatron Bridge. Synchronization, transport, and inference-specific loading would remain in downstream RL frameworks.

Describe alternatives you've considered

  • Read TE metadata or internal buffers downstream. This works today but requires each integration to understand and maintain TE-specific extraction logic.
  • Always convert to BF16 and requantize. This adds unnecessary work when compatible quantized storage already exists.

Additional context

Examples of relevant work:

Ngôn ngữ chính
Python
Star
3.6k
Fork
851
Merge trung bình
4 ngày 15 giờ
Pull request đã merge (30 ngày)
51

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của NVIDIA/TransformerEngine

Tất cả issue của NVIDIA/TransformerEngine

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.