Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar

Open Beginner friendly
#3,636 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 2 days

@sanjana658 is already working on this.

Since Oct 8, 2026.

  • #3659 by @sanjana658 — open

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
82/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python, pytorch

Research direction

Read transformer_engine/pytorch/onnx_extensions.py around onnx_cs_quantize_fp8_op and its register_fake implementation. Change the fake inverse-scale output to match the eager scalar shape, then run the proposed torch.library.opcheck FakeTensor check in a compatible TE/CUDA environment; done means eager and fake output metadata match.

Written by the indexing model from the issue text.

Description

Problem

At current main 9d4bd38678f29a29dc0989e3a679dc14f3538e6a, onnx_cs_quantize_fp8_op computes amax = tensor.abs().max() without a dimension, then scale_inv = 1 / scale. This produces a zero-dimensional float32 tensor. Its register_fake implementation instead returns torch.ones(1, ...) for that output, describing a one-dimensional tensor.

This violates the custom operator's eager/FakeTensor metadata contract. PyTorch's torch.library.opcheck(..., test_utils=("test_faketensor",)) reports:

found mismatched tensor metadata for output[1]:
Shapes torch.Size([]) and torch.Size([1]) are not equal!
Narrow reproduction and limitation

I isolated the exact upstream current-scaling function and its fake registration on Linux with PyTorch 2.11.0+cu130, CUDA hidden. I replaced only the lower tex::fp8_quantize operation with a CPU uint8 shape producer. The separately computed inverse scale does not depend on that producer's output. Across FP32/FP16/BF16 inputs with shapes (4,), (2, 3), and (2, 3, 4), all nine original metadata checks fail with the mismatch above.

Changing only the fake inverse-scale allocation to torch.ones((), dtype=torch.float32, device=tensor.device) makes all nine isolated metadata checks pass. This is a shape-contract reproduction, not a test of native TE FP8 kernels, numerical quantization, ONNX Runtime, or TensorRT export.

Suggested fix and native follow-up

Return a scalar inverse scale in the fake implementation:

return torch.empty(tensor.shape, dtype=torch.uint8, device=tensor.device), torch.ones(
    (), dtype=torch.float32, device=tensor.device
)

On a compatible installed TE/CUDA environment, a focused regression can compare eager and FakeTensor metadata using:

import torch
import transformer_engine.pytorch.onnx_extensions

x = torch.randn(16, 16, device="cuda", dtype=torch.float32)
torch.library.opcheck(torch.ops.tex.fp8_cs_quantize.default, (x,),
                     test_utils=("test_faketensor",))

The native follow-up command above is proposed, not executed in my isolated reproduction.

Dominant language
Python
Stars
3.6k
Forks
851
Avg merge
5d 1h
Merged PRs (30d)
52

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/TransformerEngine

All issues in NVIDIA/TransformerEngine

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.