Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Question] Expected behavior for blockwise FP8? Hybrid E4M3/E5M2 format & eval metrics outperforming BF16

未关闭
#2,754 2 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

维护者通常 2 天内回复

@ptrendx 已经在做这个了。

开始于 2026年3月17日。

评估

这个 Issue 还没有评估数据。

描述

question

Hi,

I am currently experimenting TE with blockwise FP8 training (DeepSeek-V3 like recipe). I am training on internal company dataset. Since I currently do not have access to SM10x-architecture devices, I am unable to utilize the MXFP8 format for these runs.

During these experiments, I observed two somewhat counter-intuitive phenomena compared to BF16 baseline. I would like to consult if these are expected behaviors under this specific training regime, or if they might indicate potential flaws in my scaling/quantization implementation.

Observation 1: Hybrid format (Fwd E4M3 + Bwd E5M2) aligns much better with BF16 than full-E4M3
When using the default e4m3 format for both forward and backward passes, the loss alignment with the BF16 baseline is suboptimal. However, switching to a hybrid approach (e4m3 for forward activations/weights, e5m2 for backward gradients) yields a much closer alignment to BF16.

Observation 2: FP8 slightly but stably outperforms BF16 in late-stage evaluation
In the final 30% of the training steps, the eval metrics for the FP8 run slightly but stably surpass the BF16 baseline.

Are these observations common and theoretically sound when using TE with fine-grained FP8 recipes? Any insights or references would be greatly appreciated.

主要语言
Python
星标
3.6k
派生
851
平均合并
5 天 1 小时
30 天内合并 PR
52

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

NVIDIA/TransformerEngine 的其他 Issue

查看 NVIDIA/TransformerEngine 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。