/v3/audio/speech ignores response_format and always returns 32-bit float WAV
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 3/5
- Tempo stimato
- 1-2 giorni
- Idoneità per principianti
- 58/100
Direzione di ricerca
Start at the /v3/audio/speech entry point and trace the T2sCalculator output used by the provided graph.pbtxt. Reproduce the curl requests for wav, pcm, mp3, and flac, then verify that supported formats have the expected sample encoding and Content-Type, while unsupported formats return HTTP 400.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Describe the bug
The OpenAI-compatible text-to-speech endpoint /v3/audio/speech ignores response_format. For wav, pcm, mp3 and flac it returns the same body: a WAV file with 32-bit IEEE float samples (format tag 3). The response is also sent with Content-Type: application/json; charset=utf-8.
The OpenAI API specifies 16-bit PCM for wav and headerless 16-bit little-endian PCM for pcm, so clients written against it cannot play the output. For example, wyoming_openai (the Wyoming bridge used by Home Assistant) cannot parse the float WAV header and plays static: https://github.com/roryeckel/wyoming_openai/issues/72
To Reproduce
- Models repository:
OpenVINO/Kokoro-82M-int8-ovpulled from Hugging Face into/models/OpenVINO/Kokoro-82M-int8-ov, graph generated withovms --configure --model_path /models/OpenVINO/Kokoro-82M-int8-ov --target_device CPU:input_stream: "HTTP_REQUEST_PAYLOAD:input" output_stream: "HTTP_RESPONSE_PAYLOAD:output" node { name: "T2sExecutor" calculator: "T2sCalculator" input_side_packet: "TTS_NODE_RESOURCES:t2s_servable" input_stream: "HTTP_REQUEST_PAYLOAD:input" output_stream: "HTTP_RESPONSE_PAYLOAD:output" node_options: { [type.googleapis.com / mediapipe.T2sCalculatorOptions]: { models_path: "/models/OpenVINO/Kokoro-82M-int8-ov" target_device: "CPU" plugin_config: '{"NUM_STREAMS":"1"}' } } } - OVMS launch command (container
openvino/model_server:latest-gpu):--rest_port 8080 --config_path /config/config.json --cache_dir /cache --log_level INFO - Client command:
for rf in wav pcm mp3 flac; do curl -s http://localhost:8080/v3/audio/speech -H 'Content-Type: application/json' \ -d "{\"model\":\"OpenVINO/Kokoro-82M-int8-ov\",\"input\":\"Hello there.\",\"voice\":\"am_adam\",\"response_format\":\"$rf\"}" \ -o "out.$rf" file -b "out.$rf" done - Every file is identical (146444 bytes):
Response headers:RIFF (little-endian) data, WAVE audio, IEEE Float, mono 24000 HzHTTP/1.1 200 OK content-length: 146444 content-type: application/json; charset=utf-8
Expected behavior
wav(the default) returns a 16-bit PCM WAV, as OpenAI does.pcmreturns headerless 16-bit little-endian PCM.- Formats OVMS cannot produce (e.g.
mp3,flac) are rejected with a 400 error instead of silently returning a WAV. - The
Content-Typematches the audio format (e.g.audio/wav).
Logs
Nothing is logged for these requests at --log_level INFO; I have not captured DEBUG logs.
Configuration
- OVMS version:
OpenVINO Model Server 2026.4.0.869b2186a, OpenVINO backend2026.4.0-22959-99c81491cc3-releases/2026/4, OpenVINO GenAI backend2026.4.0.0-3407-7ea2546852a - config.json:
{ "model_config_list": [ { "config": { "name": "OpenVINO/Kokoro-82M-int8-ov", "base_path": "/models/Kokoro-82M-int8-ov-cpu" } } ] } - Intel Core Ultra 5 125H (Meteor Lake), target device CPU
- Model repository:
/models/Kokoro-82M-int8-ov-cpu/graph.pbtxt /models/OpenVINO/Kokoro-82M-int8-ov/{config.json, openvino_config.json, openvino_model.xml, openvino_model.bin, voices/, ...} - Model: https://huggingface.co/OpenVINO/Kokoro-82M-int8-ov
- Lingua principale
- C++
- Stelle
- 932
- Fork
- 278
- Merge medio
- 2g 23h
- PR unite (30g)
- 68
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di openvinotoolkit/model_server
-
enhancement
Difficoltà 3/5 1-2 giorni Idoneità per principianti 68/100
openvinotoolkit/model_server#4609 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedForse già presa @atobiszei l’ha presa 4 giorni fa. Aperta
openvinotoolkit/model_server#4604 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking upForse già presa @atobiszei l’ha presa 4 giorni fa. Aperta
openvinotoolkit/model_server#4603 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
openvinotoolkit/model_server#4599 · 4 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 55/100
openvinotoolkit/model_server#4586 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di openvinotoolkit/model_server
Issue simili
-
`enzymexla.linalg.lu` lowering fails for a tall matrix: the permutation is built with the pivot typeAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
EnzymeAD/Enzyme-JAX#3286 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
apache/iceberg-cpp#973 ·
I maintainer di solito rispondono entro 1 giorno
-
Add c++23 mapping to nvccApertafeature request
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 86/100
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno