Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Streaming RNNT wedges after sustained zero-PCM silence; Vulkan aborts with GGML_ASSERT(ne3 == ne13)

Aperta
#48 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 3 giorni

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
48/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
cpp

Direzione di ricerca

Inizia dal percorso di streaming WebSocket /v1/realtime e da CacheAwareEncoder::encode, utilizzando il comando del server fornito e la matrice di temporizzazione zero-PCM per riprodurre il problema su CPU e Vulkan. È completato quando il parlato successivo a un silenzio zero-PCM completo e prolungato produce deltas e un final, e lo stesso schema non raggiunge più l’assertion di Vulkan segnalata.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Summary

The streaming RNNT server (Nemotron-3.5, cache-aware) wedges mid-stream after a long run of zero-PCM silence: the next real speech produces no deltas and no final until the client tears the stream down. On the Vulkan backend the same session can instead abort overnight with GGML_ASSERT(ne3 == ne13) inside CacheAwareEncoder::encode.

Reproduced in full isolation (a fresh server, a raw WebSocket client, no application code), on both the Vulkan and CPU backends, so this is in the server/streaming path, not a client bug.

Environment

  • Binary: nemo-speech 0.1.0 (manually deployed linux x86_64 vulkan build; exact source commit unknown)
  • OS/GPU: Arch Linux, kernel 7.2.3-arch1-3; Intel Integrated GPU (LNL), Mesa Vulkan 1.4.354
  • Model: nemotron-3.5 (indexed nvidia/nemotron-3.5-asr-streaming-0.6b, rev 1c8deaecc64b91f034d73e08dd8b64625eb3395d), RNNT head
  • Serve: nemo-speech serve --asr-model nemotron-3.5 --device vulkan:0|cpu --no-ui --no-warmup --asr.endpointing.enable=true --host 127.0.0.1 --port 8080

Detailed repro

Client opens GET /v1/realtime (WebSocket), sends one session.update (language=es-ES, sample_rate=16000, endpointing_ms=2500), then streams 16 kHz mono PCM16 in 160 ms frames paced in real time. The stimulus is a fixed 4.9 s Spanish clip followed by trailing silence (to trigger a .completed), then a "gap", then the same clip again.

gap contents 15s 20s 25s 60s result
wire-silent (no bytes) — 2 finals 2 finals — 2nd clip decodes fine
1 ms zero-PCM frames every 160 ms — — — 2 finals (30/45/60s) 2nd clip decodes fine
16/40/80 ms zero-PCM frames every 160 ms 2 finals 2 finals 2 finals — 2nd clip decodes fine
160 ms zero-PCM frames (full) 1 final 1 final 1 final 1 final 2nd clip produces 0 deltas, 0 finals

Notes:

  • Full 160 ms zero-PCM frames are exactly what a downstream noise gate emits while idle, so this is a real-world pattern, not synthetic.
  • input_audio_buffer.clear during the gap does not recover the stream — the wedge persists.
  • With --asr.endpointing.enable off, 60 s of full zero-PCM frames still decodes continuously (no EOU, but hundreds of deltas), so the wedge is specific to the token-silence EOU path, not the base cache-aware RNNT decoder.

Crash (Vulkan backend)

Running the same streaming pattern (real speech + long idle silences) on the Vulkan backend, the server eventually aborts:

GGML_ASSERT(ne3 == ne13) failed
/work/ggml/src/ggml-cpu/ggml-cpu.c:1270

Backtrace (excerpt, coredumpctl):

#3  ggml_abort (libggml-base.so.0)
#4  ggml_compute_forward_mul_mat (libggml-cpu.so.0)
#11 ggml_backend_sched_graph_compute_async (libggml-base.so.0)
#13 CacheAwareEncoder::EncoderBatcher::...::operator() (libnemo_speech_asr.so)
#14 CacheAwareEncoder::encode (libnemo_speech_asr.so)

Nine nemo-speech SIGABRT cores exist on this machine, all the ne3 == ne13 assert (mix of the known intermittent warmup abort and this mid-stream form). The mid-stream one occurs specifically while decoding a live stream that has sat idle in zero-PCM silence.

Hypothesis

Token-silence EOU ends a long silence run by leaving decoder/lattice state that a subsequent fire_eou + reset_utterance does not fully clear (clear also does not). The next non-silent chunk is still emitted into that wedged state, so it decodes nothing. On Vulkan the same residual state surfaces on the next CacheAwareEncoder::encode as a cache-dimensions mismatch (ne3 == ne13), aborting the process.

Lingua principale
C++
Stelle
167
Fork
37
Merge medio
7g 17h
PR unite (30g)
8

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di NVIDIA/NeMo-Speech.cpp

Tutte le issue di NVIDIA/NeMo-Speech.cpp

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.