Streaming RNNT wedges after sustained zero-PCM silence; Vulkan aborts with GGML_ASSERT(ne3 == ne13)
I maintainer di solito rispondono entro 3 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- cpp
- Ambito
- audio-video-rtc, backend
Direzione di ricerca
Inizia dal percorso di streaming WebSocket /v1/realtime e da CacheAwareEncoder::encode, utilizzando il comando del server fornito e la matrice di temporizzazione zero-PCM per riprodurre il problema su CPU e Vulkan. È completato quando il parlato successivo a un silenzio zero-PCM completo e prolungato produce deltas e un final, e lo stesso schema non raggiunge più l’assertion di Vulkan segnalata.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
The streaming RNNT server (Nemotron-3.5, cache-aware) wedges mid-stream after a long run of zero-PCM silence: the next real speech produces no deltas and no final until the client tears the stream down. On the Vulkan backend the same session can instead abort overnight with GGML_ASSERT(ne3 == ne13) inside CacheAwareEncoder::encode.
Reproduced in full isolation (a fresh server, a raw WebSocket client, no application code), on both the Vulkan and CPU backends, so this is in the server/streaming path, not a client bug.
Environment
- Binary: nemo-speech
0.1.0(manually deployedlinux x86_64 vulkanbuild; exact source commit unknown) - OS/GPU: Arch Linux, kernel
7.2.3-arch1-3; Intel Integrated GPU (LNL), Mesa Vulkan1.4.354 - Model:
nemotron-3.5(indexednvidia/nemotron-3.5-asr-streaming-0.6b, rev1c8deaecc64b91f034d73e08dd8b64625eb3395d), RNNT head - Serve:
nemo-speech serve --asr-model nemotron-3.5 --device vulkan:0|cpu --no-ui --no-warmup --asr.endpointing.enable=true --host 127.0.0.1 --port 8080
Detailed repro
Client opens GET /v1/realtime (WebSocket), sends one session.update (language=es-ES, sample_rate=16000, endpointing_ms=2500), then streams 16 kHz mono PCM16 in 160 ms frames paced in real time. The stimulus is a fixed 4.9 s Spanish clip followed by trailing silence (to trigger a .completed), then a "gap", then the same clip again.
| gap contents | 15s | 20s | 25s | 60s | result |
|---|---|---|---|---|---|
| wire-silent (no bytes) | — | 2 finals | 2 finals | — | 2nd clip decodes fine |
| 1 ms zero-PCM frames every 160 ms | — | — | — | 2 finals (30/45/60s) | 2nd clip decodes fine |
| 16/40/80 ms zero-PCM frames every 160 ms | 2 finals | 2 finals | 2 finals | — | 2nd clip decodes fine |
| 160 ms zero-PCM frames (full) | 1 final | 1 final | 1 final | 1 final | 2nd clip produces 0 deltas, 0 finals |
Notes:
- Full 160 ms zero-PCM frames are exactly what a downstream noise gate emits while idle, so this is a real-world pattern, not synthetic.
input_audio_buffer.clearduring the gap does not recover the stream — the wedge persists.- With
--asr.endpointing.enableoff, 60 s of full zero-PCM frames still decodes continuously (no EOU, but hundreds of deltas), so the wedge is specific to the token-silence EOU path, not the base cache-aware RNNT decoder.
Crash (Vulkan backend)
Running the same streaming pattern (real speech + long idle silences) on the Vulkan backend, the server eventually aborts:
GGML_ASSERT(ne3 == ne13) failed
/work/ggml/src/ggml-cpu/ggml-cpu.c:1270
Backtrace (excerpt, coredumpctl):
#3 ggml_abort (libggml-base.so.0)
#4 ggml_compute_forward_mul_mat (libggml-cpu.so.0)
#11 ggml_backend_sched_graph_compute_async (libggml-base.so.0)
#13 CacheAwareEncoder::EncoderBatcher::...::operator() (libnemo_speech_asr.so)
#14 CacheAwareEncoder::encode (libnemo_speech_asr.so)
Nine nemo-speech SIGABRT cores exist on this machine, all the ne3 == ne13 assert (mix of the known intermittent warmup abort and this mid-stream form). The mid-stream one occurs specifically while decoding a live stream that has sat idle in zero-PCM silence.
Hypothesis
Token-silence EOU ends a long silence run by leaving decoder/lattice state that a subsequent fire_eou + reset_utterance does not fully clear (clear also does not). The next non-silent chunk is still emitted into that wedged state, so it decodes nothing. On Vulkan the same residual state surfaces on the next CacheAwareEncoder::encode as a cache-dimensions mismatch (ne3 == ne13), aborting the process.
- Lingua principale
- C++
- Stelle
- 167
- Fork
- 37
- Merge medio
- 7g 17h
- PR unite (30g)
- 8
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di NVIDIA/NeMo-Speech.cpp
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 52/100
NVIDIA/NeMo-Speech.cpp#61 ·
I maintainer di solito rispondono entro 3 giorni
-
enhancement
NVIDIA/NeMo-Speech.cpp#56 · 3 commenti ·
I maintainer di solito rispondono entro 3 giorni
-
Expose confidence estimation / token probabilities (currently always 1.0)Forse già presa @pskrunner14 l’ha presa 8 giorni fa. Apertaenhancement
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
NVIDIA/NeMo-Speech.cpp#55 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 3 giorni
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 52/100
NVIDIA/NeMo-Speech.cpp#49 · 1 commento · 1 reazione ·
I maintainer di solito rispondono entro 3 giorni
-
Token-silence EOU misfires mid-sentence, hard reset corrupts transcriptForse già presa @ryanleary l’ha presa 28 giorni fa. Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 45/100
NVIDIA/NeMo-Speech.cpp#40 ·
I maintainer di solito rispondono entro 3 giorni
Tutte le issue di NVIDIA/NeMo-Speech.cpp
Issue simili
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
EsotericSoftware/spine-runtimes#3186 ·
-
An empty line splits a signature where an ordinary comment is right above an argument's HaddockAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
mrkkrp/tilia#213 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Round video messages start gray and blocky with libx264: encoder is configured for 1,000,000 fpsAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
telegramdesktop/tdesktop#31422 ·
I maintainer di solito rispondono entro 9 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
duckdb/duckdb-quack#299 ·
I maintainer di solito rispondono entro 1 giorno