Vendor-neutral reference: NVFP4 vs OCP MXFP4 block-structure parameters (request for confirmation)
Maintainer thường phản hồi trong vòng 2 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 25/100
- Loại issue
- Tài liệu
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Ít trao đổi
- Lĩnh vực
- documentation, machine-learning, performance
Hướng nghiên cứu
Bắt đầu bằng việc xem xét triển khai TransformerEngine NVFP4 và mọi unit test bao quát việc đóng gói FP4, scale, bão hòa và round trip, sau đó so sánh chúng với NVIDIA developer blog được liên kết. Issue được hoàn tất khi các maintainer xác nhận hoặc sửa kích thước block, mã hóa scale, endianness, các vector tham chiếu và hành vi ngoài phạm vi.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Hi TransformerEngine maintainers,
I am compiling a vendor-neutral numeric format catalog (84 formats, 13
families) with bit-exact conformance vectors. The catalog is open and
lives at https://github.com/gHashTag/t27. NVFP4 is on the near-term
roadmap (Track 2) but I would like to ground its row entry in
parameters confirmed by the upstream implementer rather than guessed
from blog posts. This issue is an information request, not a bug
report.
What I have so far
Based on the public NVIDIA developer blog
(https://developer.nvidia.com/blog/introducing-nvfp4-for-efficient-and-accurate-low-precision-inference/)
and code references in TransformerEngine, I have populated the
following parameter table for NVFP4 alongside its closest OCP MX
counterpart (MXFP4):
| Parameter | OCP MX MXFP4 | NVIDIA NVFP4 |
|---|---|---|
| Element layout | S1E2M1 (4 bits) | S1E2M1 (4 bits) |
| Block size | 32 elements | 16 elements |
| Scale format | E8M0 (8-bit) | FP8 E4M3 (8-bit) |
| Scale exponent bits | 8 (pure exponent) | 4 |
| Scale mantissa bits | 0 | 3 |
| Scale dynamic range | 2^-127 to 2^127 | ~2^-9 to 448 |
| Scale granularity per decade | 1 (power-of-two) | 8 (3-bit mantissa) |
| Bits/element including scale | 4 + 8/32 = 4.25 | 4 + 8/16 = 4.50 |
The element layout S1E2M1 is bit-identical between the two formats;
they diverge at the block-and-scale level. Three structural
consequences follow:
- NVFP4 resolves intra-block dynamic range 8x more finely than MXFP4
within its representable range (3-bit mantissa on FP8 E4M3 scale). - NVFP4 cannot represent per-block scales outside FP8 E4M3 range
(saturates at 448, underflows below ~2^-9) without higher-level
rescaling; MXFP4 spans a much wider scale range via E8M0. - Effective bits per element differ: 4.25 (MXFP4) vs 4.50 (NVFP4),
a 5.9% overhead delta in NVFP4 that any compression-ratio
comparison should account for.
Specific requests
If a maintainer could confirm or correct any of the following, that
would close out the row and let me publish a sister conformance pack
to the existing MXFP4 pack:
(a) Block size confirmation. Is 16 elements per block the only
supported block size, or is it a default with alternatives?
(b) Scale format confirmation. Is FP8 E4M3 (with the standard
fn saturation flag, no infinities) the canonical scale
encoding? Are there variants that use FP8 E5M2 instead?
(c) Encoding endianness. When 16 four-bit elements are packed
into 64 bits, are the first element bits in the most-significant
or least-significant nibble?
(d) Reference vectors. Does TransformerEngine ship any unit
tests with documented input/output bit-patterns that I can use
as ground-truth boundary vectors (NaN, +/-Inf-equivalent
saturation, smallest normal, smallest subnormal, denormal-block
behavior)?
(e) Round-trip behavior on out-of-range scale. When a tensor's
natural per-block scale would land outside the FP8 E4M3
representable range, is the recommended behavior (i) clamp the
scale and saturate the elements, (ii) error out, or (iii)
something else?
What I will do with confirmed answers
Open a small PR (catalog row + conformance pack) on
gHashTag/t27, with full attribution to this issue and a
cross-link back to the relevant TransformerEngine references. The
pack will follow the same shared row schema as the existing six
packs (GF16, MXFP4 element, BF16, FP8 E4M3, FP8 E5M2, E8M0 block
scale), with honest abs_error reporting (no overflow-to-Inf
masked as a match).
Background and methodology are documented in a 16-page methodology
paper (Trinity S^3 AI, 2026-06-08, file
paper3-methodology-2026-06-08-v3-trinity.pdf, SHA-256
f31f5dd243afc7b2ba4a423859a1e1dc67036c3a93affab30acc8d02f0a15eef)
that I plan to upload to arXiv this week.
Async only -- no rush, no specific deadline. If the relevant
maintainer is on vacation or sprint-locked, a one-line "ping us back
in N weeks" is a fine answer.
Thank you for the open release of NVFP4 documentation and for the
maintained NVFP4 reference implementation in TransformerEngine.
-- Dmitrii Vasilev
Trinity S^3 AI
[email protected]
GitHub: @gHashTag
- Ngôn ngữ chính
- Python
- Star
- 3.6k
- Fork
- 851
- Merge trung bình
- 5 ngày 1 giờ
- Pull request đã merge (30 ngày)
- 52
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/TransformerEngine
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on itĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA/TransformerEngine#3647 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalarCó thể đã có người làm @sanjana658 đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
NVIDIA/TransformerEngine#3636 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run itCó thể đã có người làm @yuweih205 đã nhận 32 ngày trước. Đang mởattention
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
NVIDIA/TransformerEngine#3481 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Increase MAX_TENSOR_NUMĐang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
NVIDIA/TransformerEngine#2189 · 7 bình luận · 5 reaction ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration buildsCó thể đã có người làm @ksivaman đã nhận hôm nay. Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 50/100
NVIDIA/TransformerEngine#3645 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của NVIDIA/TransformerEngine
Issue tương tự
-
HTML: <template> content is extracted as document textCó thể đã có người làm @ryanmeowy đã nhận hôm nay. Đang mởbug html
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 82/100
docling-project/docling#4714 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[BUG] Qdrant RAG client applies score_threshold to raw cosine similarity, not the 0-1 score it returnsCó thể đã có người làm @roydonsequeira đã nhận hôm nay. Đang mởbug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Host test failure in core/direct_io.zig on Linux kernel 6.17: O_DIRECT open succeeds on procfs, so the test's 'plain' fd is not plainCó thể đã có người làm Có pull request liên kết đang mở hoặc đã được merge. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
ashhart/TensorFold#536 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area/install-update comp/cli duplicate P2 python:uv sweeper:risk-compatibility type/bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 62/100
NousResearch/hermes-agent#135440 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Deepak3699/Ai_Mentor#244 ·
Maintainer thường phản hồi trong vòng 1 ngày