/chat/completions fails ("Chat template not loaded correctly") for VL-derived / omni text decoders — Python-Jinja path leaves template null; MINJA mode would fix it but is not in the 2026.2.x release
I maintainer di solito rispondono entro 1 giorno
@dkalinowski ci sta già lavorando.
Dal 25/6/2026.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Summary
On a _python_on Windows build of OVMS, /v3/chat/completions fails for a text-decoder IR extracted from a tri-modal (text+image+audio+video) qwen3_5 model:
Mediapipe execution failed. MP status - INVALID_ARGUMENT: CalculatorGraph::Run() failed:
Calculator::Process() for node "LLMExecutor" failed:
Error: Chat template not loaded correctly, so it cannot be applied
/v3/completions (raw prompt) on the same served model works perfectly, so the model, tokenizer, and inference are fine — only the chat-template application path fails. The model is loaded as an LLM (continuous-batching) pipeline, which defaults to the Python-Jinja2 template processor.
Environment
- OVMS
2026.2.1(ovms_windows_2026.2.1_python_on), GenAI backend2026.2.1.0-3123. Also reproduced on2026.2.0. - Windows, Intel Arc GPU (
targetDevice: GPU). - Model: text decoder extracted from a
qwen3_5omni model, exported to INT4 OpenVINO IR viaoptimum-cli. Standard Qwen ChatML template;<|im_start|>/<|im_end|>present in vocab;bos=None, eos=<|im_end|>(identical to a working Qwen3-14B).
Reproduction
- Serve the omni-derived text-decoder IR as an LLM continuous-batching pipeline (default
graph.pbtxt). POST /v3/completionswith a raw ChatML prompt → works, correct output.POST /v3/chat/completionswithmessages→ fails with the error above.
Root cause (traced in source)
/chat uses the embedded Python Jinja2 processor, and the template object ends up null:
src/llm/py_jinja_template_processor.cpp(~L39-40):if (templateProcessor.chatTemplate == nullptr) { output = "Error: Chat template not loaded correctly, so it cannot be applied"; return false; }src/llm/servable_initializer.cpp→loadPyTemplateProcessor(~L147+): reads
tokenizer.get_original_chat_template()then compiles it in an
ImmutableSandboxedEnvironment. For this tokenizer the load/compile does not
produce a usable template, sochatTemplatestays null and/chatfails at apply time.
Importantly, GenAI's own Tokenizer.apply_chat_template() succeeds on the exact same tokenizer (verified standalone with openvino_genai 2026.2.1.0, which renders correct ChatML). So the failure is specific to OVMS's Python-Jinja serving path, not GenAI's template engine. Upgrading GenAI alone does not fix /chat.
The mechanism to fix it already exists in main — but not in the release
main has LLMCalculatorOptions.chat_template_mode (src/llm/llm_calculator.proto):
MINJA = 0— use GenAIapply_chat_template(the path that works here). "default for VLM pipelines."JINJA = 1— Python Jinja2. "default for LLM pipelines" — i.e. the failing path for this model.
There is even an in-code TODO(dkalinow) to make MINJA the default for VLM. Setting chat_template_mode: MINJA in the graph would route through the working engine — but the 2026.2.1 release binary rejects the field:
libprotobuf ERROR ... text_format.cc: Message type "mediapipe.LLMCalculatorOptions"
has no field named "chat_template_mode".
So the option is main-only and the graph fails to load when it's added on 2026.2.1.
Requests
- Release the
chat_template_modeoption in a2026.2.x/2026.3build so users can opt VL-derived/omni LLM pipelines intoMINJA. - (Robustness) In
loadChatTemplate/loadPyTemplateProcessor, when the Python-Jinja processor leaveschatTemplate == nullptr, auto-fall-back toMINJA(GenAI's engine) instead of failing/chatoutright — GenAI already handles these templates correctly. - Consider making
MINJAthe default (or auto-selected) for LLM pipelines whose tokenizer originates from a VL/omni model (aligns with the existing VLM TODO).
Current workaround
A thin reverse proxy that applies ChatML itself and forwards to /v3/completions restores /chat/completions fully (verified, correct outputs). Happy to share if useful.
- Lingua principale
- C++
- Stelle
- 932
- Fork
- 278
- Merge medio
- 3g 5h
- PR unite (30g)
- 70
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di openvinotoolkit/model_server
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 58/100
openvinotoolkit/model_server#4613 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 3/5 1-2 giorni Idoneità per principianti 68/100
openvinotoolkit/model_server#4609 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
`/v3/models` lists every model twice when `group_name` and `--idle_unload_timeout_seconds` are combinedForse già presa @atobiszei l’ha presa 5 giorni fa. Aperta
openvinotoolkit/model_server#4604 · 1 reazione · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Idle unload never happens again if the client disconnects while a sleeping graph is waking upForse già presa @atobiszei l’ha presa 5 giorni fa. Aperta
openvinotoolkit/model_server#4603 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
openvinotoolkit/model_server#4599 · 4 commenti ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di openvinotoolkit/model_server
Issue simili
-
Broken links in the docsAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 85/100
microsoft/onnxruntime#33018 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
DataLakeFileSystemClient::ListPaths() throws JSON exception due to accessing undefined fieldsApertacustomer-reported needs-triage question
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
Azure/azure-sdk-for-cpp#7435 ·
I maintainer di solito rispondono entro 1 giorno
-
ChromieCraft Generic Confirmed World Event
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
azerothcore/azerothcore-wotlk#27882 ·
I maintainer di solito rispondono entro 1 giorno
-
MacOS build failureApertabug
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
aristocratos/btop#1874 ·
I maintainer di solito rispondono entro 1 giorno