[Feature Request] Add Gated Delta Net (GDN) support
维护者通常 2 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 25/100
- Issue 类型
- 功能
- 描述清晰度
- 需要澄清
- 活跃度
- 冷清
- 技术栈
- python
调研方向
首先定位现有的 GDN 层及其 fla.ops.gated_delta_rule 和 causal_conv1d 路径。定义原生 TE kernel 对 chunked training、推理、packed sequences、BF16 和 FP8 的支持范围,然后在 GB200 上对训练性能进行基准测试。当 Megatron Core 在不依赖这些外部依赖的情况下使用原生路径,并覆盖所请求的功能时,即视为完成。
由索引模型根据 Issue 内容生成。
描述
Is your feature request related to a problem? Please describe.
Megatron Core already has a Gated Delta Net (GDN) layer, but the current implementation depends on external Triton-based fla gated-delta-rule kernels and causal_conv1d. This means GDN does not benefit from native TE kernel performance, and users need extra dependencies.
The current training kernel path is also not performing well in practice, and the gap is noticeable on GB200.
In addition, the current path still has functional gaps: inference is not supported and packed sequences are not supported.
Describe the solution you'd like
Add native TE kernels for Gated Delta Net (GDN), with Megatron Core integration.
Ideally this would include:
- an optimized kernel path for the chunked gated delta rule
- substantially better training performance, especially on GB200
- compatibility with TE mixed-precision flows, especially BF16 and FP8
Describe alternatives you've considered
Using the current Triton-based fla.ops.gated_delta_rule + causal_conv1d implementation in Megatron Core.
- 主要语言
- Python
- 星标
- 3.6k
- 派生
- 851
- 平均合并
- 5 天 1 小时
- 30 天内合并 PR
- 52
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/TransformerEngine 的其他 Issue
-
[Bug] group_quantize_fp8_blockwise: mbarrier invalidated before other threads finish waiting on it未关闭
难度 2/5 1-3 小时 新手友好度 82/100
NVIDIA/TransformerEngine#3647 ·
维护者通常 2 天内回复
-
[PyTorch] fp8_cs_quantize fake implementation returns a vector inverse scale instead of a scalar可能已有人在做 @sanjana658 于 1 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 82/100
NVIDIA/TransformerEngine#3636 · 2 条评论 ·
维护者通常 2 天内回复
-
[Bug] Backend selection picks FA3 for training with head_dim_qk=192 / v_head_dim=128, but FA3 backward cannot run it可能已有人在做 @yuweih205 于 32 天前认领。 未关闭attention
难度 2/5 1-3 小时 新手友好度 85/100
NVIDIA/TransformerEngine#3481 · 4 条评论 ·
维护者通常 2 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 68/100
NVIDIA/TransformerEngine#2189 · 7 条评论 · 5 个 reaction ·
维护者通常 2 天内回复
-
[PyTorch] CUDA graph RNG registration floods training logs on automatic-registration builds可能已有人在做 @ksivaman 于 1 天前认领。 未关闭
难度 4/5 3-5 天 新手友好度 50/100
NVIDIA/TransformerEngine#3645 · 1 条评论 · 已指派 1 人 ·
维护者通常 2 天内回复
查看 NVIDIA/TransformerEngine 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 82/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 62/100
-
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
[BUG] Multi-day events show "Ended" while still in progress可能已有人在做 @tarunagnihotri534 今天认领。 未关闭bug
难度 2/5 1-3 小时 新手友好度 85/100
data-umbrella/du-event-board#231 · 2 条评论 ·
-
avl_automation: the generated control surface block isn't valid XML (typo in avl_out_parse.py)可能已有人在做 @brksol 今天认领。 未关闭
难度 1/5 1 小时以内 新手友好度 92/100
PX4/PX4-gazebo-models#164 ·