Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Token-silence EOU misfires mid-sentence, hard reset corrupts transcript

Open
#40 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 5 days

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
45/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp
Domain
backend, cli

Research direction

Start in src/asr/runner.cpp at CacheStreamRunner::poll_endpoint, BufferedStreamRunner::poll_endpoint, and fire_eou/finish_endpoint; run the provided nemo-speech transcribe streaming reproduction. Trace how endpoint timing and reset state reach StreamingUpdate.is_final, then add regression coverage for sub-threshold pauses, decode latency, and a correctly fired boundary while preserving continuous text.

Written by the indexing model from the issue text.

Description

Bug report + fix plan: token-silence EOU misfires mid-sentence, and the resulting reset corrupts the transcript

Summary

With asr.endpointing.enable on and no VAD model loaded (the default "token-silence"
method), the cache-aware RNNT streaming path (CacheStreamRunner) fires a mid-stream
EOU (is_final) well before any real ~800ms silence occurs, on entirely ordinary
conversational pauses. Each fire hard-resets all decoder state
(all_tokens_/transcript_/predictor recurrent state), and the next segment gets
independently punctuated and capitalized as if it were a new utterance — producing
wrong, client-visible output for a continuous sentence. Reproduced on plain CLI
nemo-speech transcribe --stream --endpointing, independent of any downstream
consumer.

Repro

nemo-speech transcribe some.wav -m nvidia/nemotron-3.5-asr-streaming-0.6b \
    --stream --endpointing --stop-history-eou-ms 700 --asr.streaming.rnnt_right_context -1

Batch (non-streaming) decode of the identical audio:

In a new work tree based on Maine, can you please have a look at how we package up
profiles and what the default profile settings are that we should

Streaming, --endpointing off:

In a new work tree based on Maine, can you please have a look at how we package up
profiles and what the default profile settings are that we should

Streaming, --endpointing on (--stop-history-eou-ms 700, matching a real client's
default EOU-auto-stop config) — deterministic across repeated runs:

In a new work tree based on Maine, can you please have a look at how? We package our
profiles and what the Default profile settings are that we should

Note both the spurious ?/capitalization at two points, and drifted word content
right at each reset boundary (up → our), consistent with the reset discarding
useful decode context, not just cosmetics.

Root cause (confirmed via targeted instrumentation, not just reading)

CacheStreamRunner::poll_endpoint (src/asr/runner.cpp, non-VAD branch):

// Token-silence: time since the decoder's last token emission frame.
const double frame_ms = model_->ms_per_enc_frame();
now_ms = static_cast<double>(total_frames_emitted_) * frame_ms;
const int64_t lef = head_ ? head_->last_emit_frame() : -1;
last_speech_ms = (lef < 0) ? 0.0 : static_cast<double>(lef + 1) * frame_ms;

now_ms is the decode clock (how many encoder frames have been processed);
last_speech_ms is the frame index of the decoder's last emitted non-blank
token
. The gap between them is compared against stop_history_eou_ms (default
800, doc'd as "trailing-silence EOU threshold"). This conflates two different
things:

  1. Genuine acoustic silence — correctly grows this gap.
  2. Decode latency — how many chunks the RNNT head needed to resolve the next
    token, which is architecturally variable (cache-aware right-context is only
    1+R encoder frames per chunk; a token/word that needs more acoustic evidence
    than one chunk provides just doesn't emit for a chunk or two, with zero acoustic
    silence involved). This also grows the same gap.

The endpointer has no way to distinguish these. Instrumented with debug logging at
every poll_endpoint call (NEMO_SPEECH_DEBUG_EOU, temporary, not included in this
patch) against the repro above:

fire now_ms last_speech_ms measured gap threshold real acoustic pause (word-timing ground truth)
#1 (how?/We) 5440 4560 880ms 700ms 720ms — inflated ~160ms
#2 (the/Default) 8320 7600 720ms 700ms no single pause ≥640ms — two separate sub-threshold pauses (≈640ms before "and", ≈480ms before "the"), with "and" simply slow to tokenize

Fire #1 is a real pause, just measured ~160ms long (about one chunk, matching the
architecture's own "~1 chunk" lag estimate in the code comment) — arguably tolerable
given the threshold's margin. Fire #2 is a clear false positive: no real silence
crossed the threshold; ordinary decode latency for one word, stacked with two
unremarkable pauses, did.

The same measurement pattern (decode-clock timestamps standing in for audio-real-time
events) also affects BufferedStreamRunner::poll_endpoint, whose own comment
documents an even larger, explicitly acknowledged lag: "the frontier leads decoding
by chunk+right_pad (~2 s)"
. This is a systemic property of the endpointer design, not
an RNNT-only edge case.

Why the reset makes it worse than a mistimed boundary

fire_eou (runner.cpp):

all_tokens.clear();
transcript.clear();
if (head)
    head->reset_utterance();

Even a correctly timed EOU still fully discards decode state today. From first
principles, nothing in this architecture requires that:

  • The encoder cache is already bounded independent of fire_eou —
    cache_filled_frames_ = std::min(cache_filled_frames_ + last_enc_T_, left_ctx)
    clamps regardless of how long the stream has run.
  • Raw audio/mel history is already trimmed independent of fire_eou —
    trim_buffers()/compact_mel_buffer() discard everything before the sliding
    window's left-context need, every step() call, not just at EOU.
  • The RNNT predictor's recurrent state (prev_token_ + LSTM cell/hidden state) is
    fixed-size regardless of sequence length — it does not grow with utterance
    duration.
  • The only structure that grows with utterance length at all is transcript_/
    all_tokens_ — plain text and a token-id vector, cheap (bytes, not audio) even
    for a very long transcript.

So there is no memory/compute justification for a blanket reset on an ordinary
conversational pause. The reset exists only to give API/CLI consumers a
client-visible "final" boundary (punctuation-worthy segment) — a reporting concern —
but is implemented as a full model-state wipe, which is a correctness concern
attached to the wrong layer.

Fix plan

Two independent fixes; either alone is a real improvement, together they close this
properly:

1. Stop measuring silence against decode-clock token timing

Replace (or gate) the token-silence signal with something anchored to actual audio,
not decoder progress. Cheapest correct option: reuse the mel features already being
computed for the encoder (no new VAD model dependency) — track "last mel frame with
above-noise-floor energy" as last_speech_ms, independent of whether the RNNT head
has committed a token for it yet. This is a small, local change to
CacheStreamRunner/BufferedStreamRunner's poll_endpoint, doesn't touch the public
API, and directly fixes fire #2's false positive (decode latency stops being
mistakable for silence) and reduces fire #1's ~160ms inflation to whatever the true
feature-extraction latency is (much smaller, bounded, and constant rather than
word-dependent).

--vad-based-eou with a real Silero model sidesteps the token-timing coupling too
(VAD already runs off raw audio), but requires bundling/loading a VAD model — the
above gets the same correctness property for free from data already in hand.

2. Make EOU a soft, reporting-only boundary, not a hard reset

Don't call head->reset_utterance()/clear all_tokens_/clear transcript_ on an
ordinary EOU fire. Instead:

  • Record a boundary marker (an index into all_tokens_/transcript_, or a
    timestamp) at fire time.
  • StreamingUpdate.is_final continues to mean "here's a good client-visible commit
    point" — compute it as the text since the last boundary, same external contract
    clients see today — but leave the encoder cache and predictor state untouched, so
    the next segment naturally continues the same sentence: correct capitalization,
    no spurious punctuation, no lost acoustic/linguistic context across the join.
  • Keep a real, separate, explicit reset path for the genuinely rare case a very
    long-running stream ever needs one (e.g. a duration/token-count budget, not an
    ordinary pause), decided by a principled criterion instead of reusing the EOU
    pause signal.

This is the more significant change (touches both runners' fire_eou/finish_endpoint
and whatever downstream code assumes a hard per-segment reset happened), but is the
one that actually matches the architecture: nothing here needs the reset for memory
reasons, so removing it isn't a tradeoff against unbounded growth — the growth is
already bounded elsewhere.

Suggested validation
  • Regression test: this repro's audio (or a synthetic equivalent — a word spanning
    2+ chunks immediately before a sub-threshold pause) must not fire EOU, and streaming
    output with --endpointing on must match batch output for continuous, natural
    speech with ordinary pauses.
  • A second synthetic test asserting that once fix #2 lands, text either side of a
    correctly-fired EOU (a real long pause) is still grammatically continuous
    (capitalization/punctuation-wise) when the underlying audio actually was one
    sentence split by a long thinking-pause.

Reported by

Found investigating a downstream streaming client's bug report of unexpected
mid-sentence capitalization. Confirmed independent of that client via the CLI repro
above; instrumentation was temporary and is not included in this patch.

Dominant language
C++
Stars
150
Forks
32
Avg merge
8d 4h
Merged PRs (30d)
9

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/NeMo-Speech.cpp

All issues in NVIDIA/NeMo-Speech.cpp

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.