Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Expose confidence estimation / token probabilities (currently always 1.0)

Open
#55 2 comments 0 reactions 1 assignee View on GitHub

Maintainers usually reply within 5 days

@pskrunner14 is already working on this.

Since Sep 30, 2026.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp
Domain
api, backend

Research direction

Start at the /v1/audio/transcriptions endpoint and trace how its confidence response field is produced, then inspect the decoder output that currently yields per-token probabilities. Use NeMo's documented entropy-based confidence_cfg as a reference for the intended behavior. Done means the API returns useful varying confidence values rather than always returning 1.0, including for incorrect, silent, or garbled audio.

Written by the indexing model from the issue text.

Description

enhancement

Description:
I'm running nemotron-3.5-asr-streaming-0.6b via nemo-speech.cpp's /v1/audio/transcriptions endpoint. The confidence field in the response is always exactly 1.0, regardless of whether the transcription is actually correct.

Concrete example: audio saying "chi è Alfonso?" (Italian) is transcribed as "che è alfonso" (wrong word) — still returns confidence: 1.0. It also returns 1.0 on pure silence and on clearly hallucinated/garbled output.

I switched to running the same model through the full NeMo Python toolkit directly (not this C++ port), enabling decoding_cfg.confidence_cfg with the entropy-based method (tsallis entropy, entropy_norm: exp) documented in NeMo's own docs. That gives real, varying per-word confidence values (e.g., 0.02–0.9 depending on the utterance), which is genuinely useful for rejecting hallucinated output — something a flat 1.0 can't do.

Is there a plan to expose this confidence estimation in nemo-speech.cpp's API? The underlying decoder already computes per-token probabilities to pick output tokens, so it seems like this would be a matter of exposing an existing internal computation rather than a fundamental limitation — similar to a related open request on transcribe.cpp (#164) asking for per-token logprobs on Parakeet TDT/Qwen3-ASR.

Happy to share the exact confidence_cfg I used if useful!

Dominant language
C++
Stars
150
Forks
32
Avg merge
8d 4h
Merged PRs (30d)
9

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA/NeMo-Speech.cpp

All issues in NVIDIA/NeMo-Speech.cpp

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.