Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

test_relpos_attention_local: local (rel_pos_local_attn) attention diverges from NeMo

未关闭
#44 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
52/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
冷清
技术栈
cpp, python

调研方向

先从 forward_local 和 test_relpos_attention_local 可执行文件开始,然后将其对转储的 pos_emb 布局和 attention window 的处理与 scripts/gen_nemo_baseline.py 进行比较。在 W=64 和 W=32 时重现 CPU f32 偏差;完成的标准是 C++ 输出会响应 W 并与 NeMo baseline 匹配,同时 chunked 和 memory 测试仍然通过。

由索引模型根据 Issue 内容生成。

描述

While building a full NeMo-baseline set to run the model-dependent test suite, I
hit a divergence in the local (Longformer) attention path that looks separate
from #39 (the streaming O(N²) fix) — filing it on its own.

Symptom

test_relpos_attention_local fails on the 110m anchor, on CPU (PARAKEET_DEVICE=cpu,
f32 GGUF), so it isn't iGPU fp16 tolerance:

[relpos_attention_local] n=47616 max|d|=3.349e+02 mean|d|=9.779e+00 (worst@47338 got=0.44526 ref=335.37750) -> FAIL

The divergence is broad (mean |d| ≈ 10, not a single element) and the worst point
is the last time frame (worst index 47338 = frame 92 of T=93, d_model=512).

It's not the --att-context-size (W) chosen for the baseline

I regenerated PARAKEET_TEST_BASELINE_LOCAL at two windows and re-ran:

W result
64 worst@47338 got=0.44526 ref=335.37750
32 worst@47338 got=0.44526 ref=356.31686

The C++ output (got) is identical across W while NeMo's ref changes — i.e.
forward_local does not respond to the window the baseline encodes. (W=128 is
correctly rejected by the test since W ≥ T.)

test_relpos_attention_local_chunked and test_relpos_attention_local_memory
pass (they use an internal brute-force reference), so the gap is specific to
the non-chunked forward_local vs the NeMo rel_pos_local_attn baseline.

Reproduce

# baseline (NeMo): local attention with a finite window over speech.wav
python scripts/gen_nemo_baseline.py \
  --model nvidia/parakeet-tdt_ctc-110m \
  --audio tests/fixtures/speech.wav \
  --att-context-size 64 --output /tmp/baseline_local.gguf

# convert the 110m anchor to f32 gguf -> PARAKEET_TEST_GGUF
PARAKEET_DEVICE=cpu \
PARAKEET_TEST_GGUF=/tmp/pk110m-f32.gguf \
PARAKEET_TEST_BASELINE_LOCAL=/tmp/baseline_local.gguf \
  ./build/tests/test_relpos_attention_local

Question

Is this a known limitation, a layout/convention mismatch between the dumped
pos_emb ([2W+1, d_model]) and what forward_local expects, or a real bug in
the non-chunked local path? Happy to dig into forward_local if it's worth a fix.

主要语言
C++
星标
786
派生
93
平均合并
9 天 19 小时
30 天内合并 PR
4

环境准备

我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

mudler/parakeet.cpp 的其他 Issue

查看 mudler/parakeet.cpp 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。