Plumb dbias request in isolation into the common backend selector
维护者通常 2 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 35/100
- Issue 类型
- 缺陷
- 描述清晰度
- 需要澄清
- 活跃度
- 冷清
- 技术栈
- python
调研方向
未识别出文件或测试。首先跟踪 Transformer Engine 路径到通用 backend selector 和 D=256 SM10x gate,然后确定 dBias 请求状态在哪里可用;完成的标准是 selector 能区分 dBias 请求,同时保持所述 gate 行为不变。
由索引模型根据 Issue 内容生成。
描述
Describe the bug
cuDNN FE can model bias input separately from dBias, but TE does not yet plumb whether dBias is requested into the common backend selector. Until that distinction is available, the D=256 SM10x gate requires no bias.
Steps/Code to reproduce bug
Please list minimal steps or code snippet for us to be able to reproduce the bug.
A helpful guide on on how to craft a minimal bug report http://matthewrocklin.com/blog/work/2018/02/28/minimal-bug-reports.
Expected behavior
A clear and concise description of what you expected to happen.
Environment overview (please complete the following information)
- Environment location: [Bare-metal, Docker, Cloud(specify cloud provider - AWS, Azure, GCP, Collab)]
- Method of Transformer Engine install: [pip install or from source]. Please specify exact commands you used to install.
- If method of install is [Docker], provide
docker pull&docker runcommands used
Environment details
If NVIDIA docker image is used you don't need to specify these.
Otherwise, please provide:
- OS version
- PyTorch version
- Python version
- Transformer Engine version
- CUDA version
- CUDNN version
Device details
- GPU model
Additional context
Add any other context about the problem here.
- 主要语言
- Python
- 星标
- 3.6k
- 派生
- 851
- 平均合并
- 4 天 15 小时
- 30 天内合并 PR
- 51
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/TransformerEngine 的其他 Issue
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar未关闭
难度 2/5 1-3 小时 新手友好度 82/100
NVIDIA/TransformerEngine#3636 ·
维护者通常 2 天内回复
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it可能已有人在做 @yuweih205 于 31 天前认领。 未关闭attention
难度 2/5 1-3 小时 新手友好度 85/100
NVIDIA/TransformerEngine#3481 · 4 条评论 ·
维护者通常 2 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 68/100
NVIDIA/TransformerEngine#2189 · 7 条评论 · 5 个 reaction ·
维护者通常 2 天内回复
-
难度 4/5 3-5 天 新手友好度 50/100
NVIDIA/TransformerEngine#3645 ·
维护者通常 2 天内回复
-
enhancement
难度 5/5 一周以上 新手友好度 35/100
NVIDIA/TransformerEngine#3644 ·
维护者通常 2 天内回复
查看 NVIDIA/TransformerEngine 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复
-
SR_SECURITY_DESCRIPTOR.fromString drops the SACL when no DACL is present可能已有人在做 @paul7436 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 85/100
equinor/fmu-sumo-uploader#302 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 75/100
modelscope/evalscope#1821 ·
维护者通常 1 天内回复
-
Sanity on ansible-core devel fails: ignore-2.23.txt references the removed import-3.9 test可能已有人在做 @yurnov 今天认领。 未关闭needs_triage
难度 1/5 1 小时以内 新手友好度 91/100
ansible-collections/kubernetes.core#1275 ·
维护者通常 1 天内回复