[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar
维护者通常 2 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 新手友好度
- 82/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 活跃
调研方向
阅读 transformer_engine/pytorch/onnx_extensions.py 中 onnx_cs_quantize_fp8_op 及其 register_fake 实现附近的代码。修改 fake 逆缩放输出,使其与 eager 标量形状一致,然后在兼容的 TE/CUDA 环境中运行建议的 torch.library.opcheck FakeTensor 检查;当 eager 和 fake 输出元数据匹配时即为完成。
由索引模型根据 Issue 内容生成。
描述
Problem
At current main 9d4bd38678f29a29dc0989e3a679dc14f3538e6a, onnx_cs_quantize_fp8_op computes amax = tensor.abs().max() without a dimension, then scale_inv = 1 / scale. This produces a zero-dimensional float32 tensor. Its register_fake implementation instead returns torch.ones(1, ...) for that output, describing a one-dimensional tensor.
This violates the custom operator's eager/FakeTensor metadata contract. PyTorch's torch.library.opcheck(..., test_utils=("test_faketensor",)) reports:
found mismatched tensor metadata for output[1]:
Shapes torch.Size([]) and torch.Size([1]) are not equal!
Narrow reproduction and limitation
I isolated the exact upstream current-scaling function and its fake registration on Linux with PyTorch 2.11.0+cu130, CUDA hidden. I replaced only the lower tex::fp8_quantize operation with a CPU uint8 shape producer. The separately computed inverse scale does not depend on that producer's output. Across FP32/FP16/BF16 inputs with shapes (4,), (2, 3), and (2, 3, 4), all nine original metadata checks fail with the mismatch above.
Changing only the fake inverse-scale allocation to torch.ones((), dtype=torch.float32, device=tensor.device) makes all nine isolated metadata checks pass. This is a shape-contract reproduction, not a test of native TE FP8 kernels, numerical quantization, ONNX Runtime, or TensorRT export.
Suggested fix and native follow-up
Return a scalar inverse scale in the fake implementation:
return torch.empty(tensor.shape, dtype=torch.uint8, device=tensor.device), torch.ones(
(), dtype=torch.float32, device=tensor.device
)
On a compatible installed TE/CUDA environment, a focused regression can compare eager and FakeTensor metadata using:
import torch
import transformer_engine.pytorch.onnx_extensions
x = torch.randn(16, 16, device="cuda", dtype=torch.float32)
torch.library.opcheck(torch.ops.tex.fp8_cs_quantize.default, (x,),
test_utils=("test_faketensor",))
The native follow-up command above is proposed, not executed in my isolated reproduction.
- 主要语言
- Python
- 星标
- 3.6k
- 派生
- 844
- 平均合并
- 4 天 55 分钟
- 30 天内合并 PR
- 51
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/TransformerEngine 的其他 Issue
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it可能已有人在做 @yuweih205 于 30 天前认领。 未关闭attention
难度 2/5 1-3 小时 新手友好度 85/100
NVIDIA/TransformerEngine#3481 · 4 条评论 ·
维护者通常 2 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 68/100
NVIDIA/TransformerEngine#2189 · 7 条评论 · 5 个 reaction ·
维护者通常 2 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 54/100
NVIDIA/TransformerEngine#3640 · 5 条评论 ·
维护者通常 2 天内回复
-
[BUG] Grouped MXFP8 quantization is not concurrency safe with multiple streams可能已有人在做 @kainzhong 于 1 天前认领。 未关闭bug
难度 4/5 3-5 天 新手友好度 25/100
NVIDIA/TransformerEngine#3630 ·
维护者通常 2 天内回复
-
[bug] NVFP4 + `torch.compile`: errors with 3D input可能已有人在做 @pggPL 于 2 天前认领。 未关闭bug
难度 2/5 1-3 小时 新手友好度 25/100
NVIDIA/TransformerEngine#3626 ·
维护者通常 2 天内回复
查看 NVIDIA/TransformerEngine 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 68/100
-
Task
难度 2/5 1-3 小时 新手友好度 65/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 86/100
war-and-code/dircue#200 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 87/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 1 天内回复