Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Expose confidence estimation / token probabilities (currently always 1.0)

Abierto
#55 2 comentarios 0 reacciones 1 asignado Ver en GitHub

Los mantenedores suelen responder en 5 días

@pskrunner14 ya está trabajando en esto.

Desde el 30/9/2026.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
48/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
cpp
Área
api, backend

Línea de trabajo

Start at the /v1/audio/transcriptions endpoint and trace how its confidence response field is produced, then inspect the decoder output that currently yields per-token probabilities. Use NeMo's documented entropy-based confidence_cfg as a reference for the intended behavior. Done means the API returns useful varying confidence values rather than always returning 1.0, including for incorrect, silent, or garbled audio.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

enhancement

Description:
I'm running nemotron-3.5-asr-streaming-0.6b via nemo-speech.cpp's /v1/audio/transcriptions endpoint. The confidence field in the response is always exactly 1.0, regardless of whether the transcription is actually correct.

Concrete example: audio saying "chi è Alfonso?" (Italian) is transcribed as "che è alfonso" (wrong word) — still returns confidence: 1.0. It also returns 1.0 on pure silence and on clearly hallucinated/garbled output.

I switched to running the same model through the full NeMo Python toolkit directly (not this C++ port), enabling decoding_cfg.confidence_cfg with the entropy-based method (tsallis entropy, entropy_norm: exp) documented in NeMo's own docs. That gives real, varying per-word confidence values (e.g., 0.02–0.9 depending on the utterance), which is genuinely useful for rejecting hallucinated output — something a flat 1.0 can't do.

Is there a plan to expose this confidence estimation in nemo-speech.cpp's API? The underlying decoder already computes per-token probabilities to pick output tokens, so it seems like this would be a matter of exposing an existing internal computation rather than a fundamental limitation — similar to a related open request on transcribe.cpp (#164) asking for per-token logprobs on Parakeet TDT/Qwen3-ASR.

Happy to share the exact confidence_cfg I used if useful!

Lenguaje dominante
C++
Estrellas
150
Forks
32
Merge medio
8 d 4 h
PR fusionados (30 d)
9

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de NVIDIA/NeMo-Speech.cpp

Todos los issues de NVIDIA/NeMo-Speech.cpp

Issues similares

Más issues de C++

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.