MagpieTTS long-form: a 4–5 character sentence fails the whole request ("longform history context cache is too short")
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 52/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- cpp
- Ambito
- audio-video-rtc, backend
Direzione di ricerca
Start by locating the MagpieTTS long-form synthesis path and the history-context cache check that reports “need 20 token(s)”; the issue does not name source files or tests. Reproduce with the provided command and compare long-form behavior for very short sentences versus the workaround. Done means short sentences no longer fail the request, while long-form output is not silently truncated.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Version: nemo-speech 0.1.0 and 0.2.0 (nemo-speech-0.{1,2}.0-linux-x86_64-cuda
release tarballs; the same input fails identically on both; the nightly was not tried), model nvidia/magpie_tts_multilingual_357m@452ef560f972
(v2602.f16.gguf), codec nvidia/nemo-nano-codec-22khz-1.89kbps-21.5fps.
RTX 3090, driver 570.153.02, Linux 6.8.
What happens: when the input is long enough for long-form mode to engage,
a single very short sentence anywhere in it makes synthesis fail:
[nemo-speech] synthesize session started
longform history context cache is too short: need 20 token(s), have 18
nemo-speech synthesize: MagpieTTS synthesis failed
[nemo-speech] synthesize session failed (exit code 1)
nemo-speech serve returns HTTP 500 "MagpieTTS synthesis failed" for the same
input. It is deterministic: the same text fails every time, with any voice.
Repro: a filler sentence repeated four times, then one short sentence,
then the filler four times again (~900 characters):
F="The service kept running normally while the operators reviewed the dashboards and the alert history in detail."
F4="$F $F $F $F"
echo "$F4 Okay. $F4" > t.txt
nemo-speech synthesize -i t.txt -o o.wav --force --seed 1
# -> exit 1, "need 20 token(s), have 18"
Which sentences fail: the same template, with only the middle sentence
changed:
| middle sentence | chars | result |
|---|---|---|
Why? |
4 | fails, "have 12" |
Yes. |
4 | fails, "have 12" |
Okay. |
5 | fails, "have 18" |
No way. |
7 | ok |
Why not? |
8 | ok |
| 12 other sentences of 11–20 chars | all ok |
The same short sentence passes in a short input, where long-form mode does not
engage. In a real 48-chunk text (each chunk ≤ 800 chars), 47 rendered and one
failed. It contained "Why? It ran out of disk space. Why? Logs were not
rotated. Why? …".
--tts.longform options: on behaves like auto (fails). off avoids
the error but truncates silently: an 800-char chunk that gives 53 s of audio
in long-form came back as 23 s, and 31 s with --steps 3000.
Expected: a short sentence is merged with its neighbour, or padded, so that
the history context reaches the minimum. It should not fail the request.
Workaround on our side: before sending text to Magpie, the client attaches
every sentence shorter than 8 characters to the previous sentence (or to the
next one when it comes first): "down. Why?" becomes "down, why?". With that,
the 48-chunk text renders 48 of 48.
- Lingua principale
- C++
- Stelle
- 150
- Fork
- 32
- Merge medio
- 7g 17h
- PR unite (30g)
- 8
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di NVIDIA/NeMo-Speech.cpp
-
enhancement
NVIDIA/NeMo-Speech.cpp#56 · 3 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Expose confidence estimation / token probabilities (currently always 1.0)Forse già presa @pskrunner14 l’ha presa 6 giorni fa. Apertaenhancement
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
NVIDIA/NeMo-Speech.cpp#55 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 52/100
NVIDIA/NeMo-Speech.cpp#49 · 1 commento · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Streaming RNNT wedges after sustained zero-PCM silence; Vulkan aborts with GGML_ASSERT(ne3 == ne13)Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
NVIDIA/NeMo-Speech.cpp#48 ·
I maintainer di solito rispondono entro 1 giorno
-
Token-silence EOU misfires mid-sentence, hard reset corrupts transcriptForse già presa @ryanleary l’ha presa 26 giorni fa. Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 45/100
NVIDIA/NeMo-Speech.cpp#40 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di NVIDIA/NeMo-Speech.cpp
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
godotengine/godot#124252 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 Meno di un'ora Idoneità per principianti 66/100
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
I maintainer di solito rispondono entro 1 giorno
-
lldb
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
llvm/llvm-project#229592 · 11 commenti ·
I maintainer di solito rispondono entro 1 giorno