/v3/audio/speech ignores response_format and always returns 32-bit float WAV
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 3/5
- Thời gian dự kiến
- 1-2 ngày
- Mức phù hợp với người mới
- 58/100
Hướng nghiên cứu
Start at the /v3/audio/speech entry point and trace the T2sCalculator output used by the provided graph.pbtxt. Reproduce the curl requests for wav, pcm, mp3, and flac, then verify that supported formats have the expected sample encoding and Content-Type, while unsupported formats return HTTP 400.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the bug
The OpenAI-compatible text-to-speech endpoint /v3/audio/speech ignores response_format. For wav, pcm, mp3 and flac it returns the same body: a WAV file with 32-bit IEEE float samples (format tag 3). The response is also sent with Content-Type: application/json; charset=utf-8.
The OpenAI API specifies 16-bit PCM for wav and headerless 16-bit little-endian PCM for pcm, so clients written against it cannot play the output. For example, wyoming_openai (the Wyoming bridge used by Home Assistant) cannot parse the float WAV header and plays static: https://github.com/roryeckel/wyoming_openai/issues/72
To Reproduce
- Models repository:
OpenVINO/Kokoro-82M-int8-ovpulled from Hugging Face into/models/OpenVINO/Kokoro-82M-int8-ov, graph generated withovms --configure --model_path /models/OpenVINO/Kokoro-82M-int8-ov --target_device CPU:input_stream: "HTTP_REQUEST_PAYLOAD:input" output_stream: "HTTP_RESPONSE_PAYLOAD:output" node { name: "T2sExecutor" calculator: "T2sCalculator" input_side_packet: "TTS_NODE_RESOURCES:t2s_servable" input_stream: "HTTP_REQUEST_PAYLOAD:input" output_stream: "HTTP_RESPONSE_PAYLOAD:output" node_options: { [type.googleapis.com / mediapipe.T2sCalculatorOptions]: { models_path: "/models/OpenVINO/Kokoro-82M-int8-ov" target_device: "CPU" plugin_config: '{"NUM_STREAMS":"1"}' } } } - OVMS launch command (container
openvino/model_server:latest-gpu):--rest_port 8080 --config_path /config/config.json --cache_dir /cache --log_level INFO - Client command:
for rf in wav pcm mp3 flac; do curl -s http://localhost:8080/v3/audio/speech -H 'Content-Type: application/json' \ -d "{\"model\":\"OpenVINO/Kokoro-82M-int8-ov\",\"input\":\"Hello there.\",\"voice\":\"am_adam\",\"response_format\":\"$rf\"}" \ -o "out.$rf" file -b "out.$rf" done - Every file is identical (146444 bytes):
Response headers:RIFF (little-endian) data, WAVE audio, IEEE Float, mono 24000 HzHTTP/1.1 200 OK content-length: 146444 content-type: application/json; charset=utf-8
Expected behavior
wav(the default) returns a 16-bit PCM WAV, as OpenAI does.pcmreturns headerless 16-bit little-endian PCM.- Formats OVMS cannot produce (e.g.
mp3,flac) are rejected with a 400 error instead of silently returning a WAV. - The
Content-Typematches the audio format (e.g.audio/wav).
Logs
Nothing is logged for these requests at --log_level INFO; I have not captured DEBUG logs.
Configuration
- OVMS version:
OpenVINO Model Server 2026.4.0.869b2186a, OpenVINO backend2026.4.0-22959-99c81491cc3-releases/2026/4, OpenVINO GenAI backend2026.4.0.0-3407-7ea2546852a - config.json:
{ "model_config_list": [ { "config": { "name": "OpenVINO/Kokoro-82M-int8-ov", "base_path": "/models/Kokoro-82M-int8-ov-cpu" } } ] } - Intel Core Ultra 5 125H (Meteor Lake), target device CPU
- Model repository:
/models/Kokoro-82M-int8-ov-cpu/graph.pbtxt /models/OpenVINO/Kokoro-82M-int8-ov/{config.json, openvino_config.json, openvino_model.xml, openvino_model.bin, voices/, ...} - Model: https://huggingface.co/OpenVINO/Kokoro-82M-int8-ov
- Ngôn ngữ chính
- C++
- Star
- 932
- Fork
- 278
- Merge trung bình
- 2 ngày 23 giờ
- Pull request đã merge (30 ngày)
- 68
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của openvinotoolkit/model_server
-
enhancement
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 68/100
openvinotoolkit/model_server#4609 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedCó thể đã có người làm @atobiszei đã nhận 4 ngày trước. Đang mở
openvinotoolkit/model_server#4604 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking upCó thể đã có người làm @atobiszei đã nhận 4 ngày trước. Đang mở
openvinotoolkit/model_server#4603 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
openvinotoolkit/model_server#4599 · 4 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 55/100
openvinotoolkit/model_server#4586 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của openvinotoolkit/model_server
Issue tương tự
-
`enzymexla.linalg.lu` lowering fails for a tall matrix: the permutation is built with the pivot typeĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
EnzymeAD/Enzyme-JAX#3286 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
apache/iceberg-cpp#973 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Add c++23 mapping to nvccĐang mởfeature request
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 86/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày