Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug] Float8CurrentScaling / Float8BlockScaling: fp8_quant_* QParams ignore use_power_2_scales / use_f32_scales passed to the constructor

Open
#3,593 2 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 2 days

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
75/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

Read the Float8CurrentScaling and Float8BlockScaling definitions in transformer_engine/common/recipe/init.py, then compare their QParams construction with NVFP4BlockScaling.post_init. Reproduce on CPU using the provided constructor examples with environment defaults set to 0. Done when explicit constructor values are reflected in the corresponding fp8_quant_* QParams, while omitted arguments still use environment-variable defaults.

Written by the indexing model from the issue text.

Description

bug

If use_power_2_scales (Float8CurrentScaling) or use_f32_scales (Float8BlockScaling) is passed to the constructor with a value different from its environment-variable default, the recipe reports the new value but its fp8_quant_* QParams retain the old value:

Float8CurrentScaling(use_power_2_scales=True)
use_power_2_scales = True
fp8_quant_fwd_inp.power_2_scale = False # expected True

Float8BlockScaling(use_f32_scales=True)
use_f32_scales = True
fp8_quant_fwd_inp.power_2_scale = True # expected False

Nothing raises an error, so the recipe configuration and the associated QParams can silently disagree.

Root cause (as far as I can tell)

In transformer_engine/common/recipe/init.py, fp8_quant_fwd_inp, fp8_quant_fwd_weight, and fp8_quant_bwd_grad are constructed in the class body using the class-level use_power_2_scales / use_f32_scales values.

They have no type annotation, so they are not dataclass fields. The QParams are therefore constructed once when the class is defined/imported. Passing an explicit value to the constructor changes the instance attribute, but does not rebuild the existing QParams.

NVFP4BlockScaling in the same file constructs its fp4_quant_* QParams in post_init, so those QParams are built per instance.

To Reproduce

NVTE_FP8_CURRENT_SCALING_POWER_2_SCALES=0

NVTE_FP8_BLOCK_SCALING_FP32_SCALES=0

python repro.py

from transformer_engine.common.recipe import (
Float8CurrentScaling,
Float8BlockScaling,
)

r = Float8CurrentScaling(use_power_2_scales=True)
print(r.use_power_2_scales, r.fp8_quant_fwd_inp.power_2_scale)

True False

r = Float8BlockScaling(use_f32_scales=True)
print(r.use_f32_scales, r.fp8_quant_fwd_inp.power_2_scale)

True True

This reproduces on CPU; no GPU or CUDA execution is required.

Expected behavior

The fp8_quant_* QParams should be constructed per instance from the effective recipe configuration, similar to NVFP4BlockScaling.

For example:

r = Float8CurrentScaling(use_power_2_scales=True)
assert r.use_power_2_scales is True
assert r.fp8_quant_fwd_inp.power_2_scale is True

r = Float8BlockScaling(use_f32_scales=True)
assert r.use_f32_scales is True
assert r.fp8_quant_fwd_inp.power_2_scale is False

The environment variables should still provide the default values when the corresponding constructor arguments are not passed.

Happy to work on this once assigned. Thanks

Dominant language
Python
Stars
3.6k
Forks
844
Avg merge
4d 20h
Merged PRs (30d)
60

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/TransformerEngine

All issues in NVIDIA/TransformerEngine

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.