TDT beam search fails above ~40-90s of audio: "zero-duration expansion did not reduce score" (greedy on the same audio is fine)
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 58/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- cpp
调研方向
使用提供的 16 kHz 单声道 PCM 用例,通过 parakeet_capi_transcribe_pcm_nbest_json 重现该失败,然后将其与 greedy 入口点和截断音频进行比较。跟踪 TDT beam-search 路径中围绕所报告不变量“zero-duration expansion did not reduce score”的部分,以及 tdt-0.6b-v3 模型的行为。完成标准是:长输入的 N-best 解码不再中止,同时现有的 greedy 和较短输入行为保持不变。
由索引模型根据 Issue 内容生成。
描述
parakeet_capi_transcribe_pcm_nbest_json aborts on longer inputs with
tdt_beam_search: zero-duration expansion did not reduce score
while greedy decoding of the same audio with the same model succeeds. It reads like an internal invariant that long inputs violate, rather than a bad-input case.
Environment
- parakeet.cpp v0.5.0, released
lib-macos-metal-arm64bundle (ABI 6) - macOS arm64, Metal backend
- Model:
mudler/parakeet-cpp-gguf→tdt-0.6b-v3-q4_k.gguf - Audio: 16 kHz mono float PCM, AMI meeting recordings (~100s each)
What fails
22 of 23 AMI clips (~100s each) fail. One succeeded, and only at beam 4.
It is not a beam-width interaction — beam 1 fails identically to beam 8:
beam_size=1 FAIL beam_size=2 FAIL beam_size=4 FAIL beam_size=8 FAIL
What succeeds on the identical audio
parakeet_capi_transcribe_pcm(greedy) — 242 words, no errorparakeet_capi_transcribe_pcm_batch_json(greedy + timestamps) — 242 words, 242 word records- The same N-best call on the first 30s or 40s of that same file
So the encoder, the model and the audio are all fine; it is specific to the beam search path at length.
Threshold varies with content, not a fixed limit
Truncating each file to N seconds and decoding at beam 4:
| file | length | 30s | 40s | 50s | 60s | 70s | 80s | 90s |
|---|---|---|---|---|---|---|---|---|
| ES2004a_FEE013 | 100.6s | ok | ok | ok | fail | fail | fail | fail |
| EN2002c_MEE071 | 100.0s | ok | ok | ok | ok | ok | ok | fail |
| ES2004b_MEO015 | 105.7s | ok | ok | fail | fail | fail | fail | fail |
All three pass at 30s and 40s and fail before 100s, at different points — consistent with something accumulating over frames rather than a hard cap.
Possibly model-specific
The same 100.6s clip decodes fine at beam 2, 4 and 8 with tdt_ctc-1.1b-q4_k.gguf. Only tdt-0.6b-v3 failed here, so it may be an interaction between that checkpoint's duration predictions and the expansion check.
Minimal reproduction
# ctypes against libparakeet.dylib from the v0.5.0 release
ptr = lib.parakeet_capi_transcribe_pcm_nbest_json(
ctx, samples_p, len(samples), 16000, 4, 1, 1, None)
# ptr is NULL; parakeet_capi_last_error(ctx) reports the message above.
# Truncating `samples` to 40s makes the same call succeed.
Possibly related to #55 (error above 5 min), though the threshold here is far lower and the message differs.
Not blocking for us — we moved to transcribe_pcm_batch_json, which gives greedy output with per-word timestamps and no beam search. Reporting because the failure is silent-until-it-isn't for anyone relying on N-best over long audio.
- 主要语言
- C++
- 星标
- 786
- 派生
- 93
- 平均合并
- 9 天 19 小时
- 30 天内合并 PR
- 4
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
mudler/parakeet.cpp 的其他 Issue
-
难度 4/5 3-5 天 新手友好度 55/100
mudler/parakeet.cpp#68 ·
-
难度 3/5 1-2 天 新手友好度 58/100
mudler/parakeet.cpp#62 ·
-
难度 3/5 1-2 天 新手友好度 48/100
mudler/parakeet.cpp#60 ·
-
难度 4/5 3-5 天 新手友好度 48/100
mudler/parakeet.cpp#59 · 3 条评论 ·
-
难度 4/5 3-5 天 新手友好度 45/100
mudler/parakeet.cpp#55 · 1 条评论 ·
查看 mudler/parakeet.cpp 的全部 Issue
相似的 Issue
-
`enzymexla.linalg.lu` lowering fails for a tall matrix: the permutation is built with the pivot type未关闭
难度 2/5 1-3 小时 新手友好度 78/100
EnzymeAD/Enzyme-JAX#3286 ·
维护者通常 1 天内回复
-
bug
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
apache/iceberg-cpp#973 ·
维护者通常 1 天内回复
-
feature request
难度 1/5 1 小时以内 新手友好度 86/100
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 1 天内回复