Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Feature Request] Add Gated Delta Net (GDN) support

已关闭
#2,884 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 2 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
25/100
Issue 类型
功能
描述清晰度
需要澄清
活跃度
冷清
技术栈
python

调研方向

首先定位现有的 GDN 层及其 fla.ops.gated_delta_rule 和 causal_conv1d 路径。定义原生 TE kernel 对 chunked training、推理、packed sequences、BF16 和 FP8 的支持范围,然后在 GB200 上对训练性能进行基准测试。当 Megatron Core 在不依赖这些外部依赖的情况下使用原生路径,并覆盖所请求的功能时,即视为完成。

由索引模型根据 Issue 内容生成。

描述

Is your feature request related to a problem? Please describe.

Megatron Core already has a Gated Delta Net (GDN) layer, but the current implementation depends on external Triton-based fla gated-delta-rule kernels and causal_conv1d. This means GDN does not benefit from native TE kernel performance, and users need extra dependencies.

The current training kernel path is also not performing well in practice, and the gap is noticeable on GB200.

In addition, the current path still has functional gaps: inference is not supported and packed sequences are not supported.

Describe the solution you'd like

Add native TE kernels for Gated Delta Net (GDN), with Megatron Core integration.

Ideally this would include:

  • an optimized kernel path for the chunked gated delta rule
  • substantially better training performance, especially on GB200
  • compatibility with TE mixed-precision flows, especially BF16 and FP8

Describe alternatives you've considered

Using the current Triton-based fla.ops.gated_delta_rule + causal_conv1d implementation in Megatron Core.

主要语言
Python
星标
3.6k
派生
851
平均合并
5 天 1 小时
30 天内合并 PR
52

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

NVIDIA/TransformerEngine 的其他 Issue

查看 NVIDIA/TransformerEngine 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。