Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

`test_comm_gemm_overlap.py::test_multi_layer_with_overlap_bf16` fails on A100

未关闭
#3,097 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 2 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
48/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
python, pytorch

调研方向

从 tests/pytorch/distributed/run_layer_with_overlap.py 和 test_comm_gemm_overlap.py::test_multi_layer_with_overlap_bf16 用例开始。运行提供的四 GPU 命令,然后在跟踪 overlap 路径的同时,将失败的 1024 token、两层配置与通过的 256 token 或单层用例进行比较。完成标准是识别并修复数值差异,同时不削弱检查,并且回归测试在报告的 A100 配置上通过。

由索引模型根据 Issue 内容生成。

描述

bug

Summary

On 4x A100, test_multi_layer_with_overlap_bf16[ TransformerLayer - BULK DGRAD/WGRAD - 2 layers - BF16 -False] fails its numerical check deterministically:

[rank0] NUMERICAL CHECK FAILED: layers.1.self_attention.layernorm_qkv.bias.grad not close
enough at index 771 with 0.1171875 vs 0.0703125 | rel. error = 0.6666666666666666
(tol = 0.025) | abs. error = 0.046875 (tol = 0.00125)

This does appear to be an precision problem, the same configuration passes with --seq-length=256 or --num-layers=1. I tried also to force NVTE_FUSED_ATTN=0 in addition to NVTE_FLASH_ATTN=0 but I got the same error

Environment

  • 4x A100 64GB (sm80), single node, NVLink (UB_SKIPMC=1 path, no CUDA Multicast)
  • TE @ 720ec27e (current main at time of writing), built with NVTE_CUDA_ARCHS=80
  • torch 2.12.0+cu126, cuDNN 9.10.2.21, flash-attn 2.8.3, CUDA runtime 12.6
  • driver: 535.274.02

Reproduction

UB_SKIPMC=1 NVTE_FLASH_ATTN=0 PYTORCH_JIT=0 NVTE_TORCH_COMPILE=0 NVTE_ALLOW_NONDETERMINISTIC_ALGO=0 \
torchrun --nproc_per_node=4 tests/pytorch/distributed/run_layer_with_overlap.py \
  --seed=42 --seq-length=1024 --batch-size=2 --num-heads=32 --head-dim=48 \
  --layer-type=TransformerLayer --num-layers=2

Suggested resolution

I would simply reduce seq-len. Let me know if youd welcom a PR for this.

主要语言
Python
星标
3.6k
派生
851
平均合并
5 天 1 小时
30 天内合并 PR
52

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

NVIDIA/TransformerEngine 的其他 Issue

查看 NVIDIA/TransformerEngine 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。