MagpieTTS long-form: a 4–5 character sentence fails the whole request ("longform history context cache is too short")
Maintainers usually reply within 3 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 52/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- cpp
- Domain
- audio-video-rtc, backend
Research direction
Start by locating the MagpieTTS long-form synthesis path and the history-context cache check that reports “need 20 token(s)”; the issue does not name source files or tests. Reproduce with the provided command and compare long-form behavior for very short sentences versus the workaround. Done means short sentences no longer fail the request, while long-form output is not silently truncated.
Written by the indexing model from the issue text.
Description
Version: nemo-speech 0.1.0 and 0.2.0 (nemo-speech-0.{1,2}.0-linux-x86_64-cuda
release tarballs; the same input fails identically on both; the nightly was not tried), model nvidia/magpie_tts_multilingual_357m@452ef560f972
(v2602.f16.gguf), codec nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps.
RTX 3090, driver 570.153.02, Linux 6.8.
What happens: when the input is long enough for long-form mode to engage,
a single very short sentence anywhere in it makes synthesis fail:
[nemo-speech] synthesize session started
longform history context cache is too short: need 20 token(s), have 18
nemo-speech synthesize: MagpieTTS synthesis failed
[nemo-speech] synthesize session failed (exit code 1)
nemo-speech serve returns HTTP 500 "MagpieTTS synthesis failed" for the same
input. It is deterministic: the same text fails every time, with any voice.
Repro: a filler sentence repeated four times, then one short sentence,
then the filler four times again (~900 characters):
F="The service kept running normally while the operators reviewed the dashboards and the alert history in detail."
F4="$F $F $F $F"
echo "$F4 Okay. $F4" > t.txt
nemo-speech synthesize -i t.txt -o o.wav --force --seed 1
# -> exit 1, "need 20 token(s), have 18"
Which sentences fail: the same template, with only the middle sentence
changed:
| middle sentence | chars | result |
|---|---|---|
Why? |
4 | fails, "have 12" |
Yes. |
4 | fails, "have 12" |
Okay. |
5 | fails, "have 18" |
No way. |
7 | ok |
Why not? |
8 | ok |
| 12 other sentences of 11–20 chars | all ok |
The same short sentence passes in a short input, where long-form mode does not
engage. In a real 48-chunk text (each chunk ≤ 800 chars), 47 rendered and one
failed. It contained "Why? It ran out of disk space. Why? Logs were not
rotated. Why? …".
--tts.longform options: on behaves like auto (fails). off avoids
the error but truncates silently: an 800-char chunk that gives 53 s of audio
in long-form came back as 23 s, and 31 s with --steps 3000.
Expected: a short sentence is merged with its neighbour, or padded, so that
the history context reaches the minimum. It should not fail the request.
Workaround on our side: before sending text to Magpie, the client attaches
every sentence shorter than 8 characters to the previous sentence (or to the
next one when it comes first): "down. Why?" becomes "down, why?". With that,
the 48-chunk text renders 48 of 48.
- Dominant language
- C++
- Stars
- 150
- Forks
- 32
- Avg merge
- 7d 17h
- Merged PRs (30d)
- 8
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/NeMo-Speech.cpp
-
enhancement
NVIDIA/NeMo-Speech.cpp#56 · 3 comments ·
Maintainers usually reply within 3 days
-
Expose confidence estimation / token probabilities (currently always 1.0)Possibly taken @pskrunner14 claimed this 7 days ago. Openenhancement
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/NeMo-Speech.cpp#55 · 2 comments · 1 assignee ·
Maintainers usually reply within 3 days
-
Difficulty 4/5 3-5 days Newbie friendliness 52/100
NVIDIA/NeMo-Speech.cpp#49 · 1 comment · 1 reaction ·
Maintainers usually reply within 3 days
-
Streaming RNNT wedges after sustained zero-PCM silence; Vulkan aborts with GGML_ASSERT(ne3 == ne13)Open
Difficulty 4/5 3-5 days Newbie friendliness 48/100
NVIDIA/NeMo-Speech.cpp#48 ·
Maintainers usually reply within 3 days
-
Token-silence EOU misfires mid-sentence, hard reset corrupts transcriptPossibly taken @ryanleary claimed this 27 days ago. Open
Difficulty 5/5 Over a week Newbie friendliness 45/100
NVIDIA/NeMo-Speech.cpp#40 ·
Maintainers usually reply within 3 days
All issues in NVIDIA/NeMo-Speech.cpp
Similar issues
-
needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
flashinfer-ai/flashinfer#6212 ·
Maintainers usually reply within 1 day
-
bug graphics
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
FlaxEngine/FlaxEngine#4295 · 2 comments ·
Maintainers usually reply within 2 days
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Algorithmiq/monoprop#390 ·
Maintainers usually reply within 1 day
-
docs
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
8-membered-ring atrop stereo lost in 2026.09.1Possibly taken A pull request linked to this issue is open or already merged. Openbug
Difficulty 2/5 Half a day Newbie friendliness 86/100
Maintainers usually reply within 2 days