Streaming: all deltas labeled 'reasoning' when the rendered prompt contains a (pre-closed) thinking marker — non-streaming parses the same output correctly
I maintainer di solito rispondono entro 3 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 3/5
- Tempo stimato
- 1-2 giorni
- Idoneità per principianti
- 52/100
Direzione di ricerca
L’issue non indica alcun file sorgente né alcun test; riproduci il comportamento con lo YAML del modello fornito e una richiesta curl in streaming, quindi individua il parser di streaming e confronta la sua classificazione con quella del percorso non-streaming. Il lavoro è completato quando l’output generato da un canale di thinking già chiuso produce delta di contenuto invece di delta di ragionamento, mentre l’output di thinking ordinario continua a essere classificato correttamente.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
LocalAI version: v4.7.1 (b224c96), localai/localai:latest-gpu-vulkan (docker), llama-cpp backend
Model: Gemma-4-12B-it (GGUF, Q8_0)
Summary
Gemma-4's official chat template disables thinking by ending the generation prompt with a pre-closed thought channel:
<|turn>model
<|channel>thought
<channel|>
With the channel already closed, the model emits plain answer text only — no thinking tokens at all. The non-streaming endpoint handles this correctly (message.content filled, no reasoning field). But in streaming mode, every single delta is emitted as delta.reasoning and delta.content stays null for the entire stream. The stream parser apparently enters the reasoning state because the prompt contains <|channel>thought, without checking that it is immediately closed — and never leaves that state, since the model (correctly) never emits a close marker itself.
Any correctly-templated "thinking disabled" Gemma-4 request triggers this, so OpenAI-compatible clients that read delta.content receive a completely empty stream while the server generates a full answer.
Reproduction
Model YAML (explicit template mirroring the GGUF's embedded jinja for the thinking-off case):
name: gemma-4-12b-it
backend: llama-cpp
parameters:
model: gemma-4-12b-it-Q8_0.gguf
gpu_layers: 99
stopwords:
- "<turn|>"
template:
chat_message: |-
<|turn>{{if eq .RoleName "assistant"}}model{{else}}{{.RoleName}}{{end}}
{{ if .Content }}{{.Content}}{{ end }}<turn|>
chat: |-
{{.Input -}}
<|turn>model
<|channel>thought
<channel|>
Non-streaming — correct:
curl -s http://localhost:8080/v1/chat/completions -H 'Content-Type: application/json' \
-d '{"model":"gemma-4-12b-it","max_tokens":40,"messages":[{"role":"user","content":"Say READY."}]}'
# → "message":{"role":"assistant","content":"READY. …"} (no reasoning field, finish_reason stop)
Streaming — broken (same request + "stream": true):
"delta":{"role":"assistant","content":null}
"delta":{"content":null,"reasoning":"READY"}
"delta":{"content":null,"reasoning":"."}
... ← 100 % of tokens arrive as reasoning; not a single content delta
Expected
The stream parser should only enter the reasoning state when the model emits a thinking-open marker (or should recognize that the marker in the prompt is immediately followed by its close <channel|>). Streaming and non-streaming should classify identical output identically.
Workarounds tried (no effect on the stream labeling)
- request level:
chat_template_kwargs: {enable_thinking: false},thinking: false,reasoning_format: "none",disable_thinking: true - model YAML:
disable_thinking: true,thinking_start_tokens: ["<NEVER_EMITTED>"],logit_bias(appears to be ignored on the llama-cpp path) - Removing the pre-close from the template is not viable: Gemma-4 then opens the thought channel itself (that pre-close is the model's official "thinking off" mechanism).
- Lingua principale
- Go
- Stelle
- 49.2k
- Fork
- 4.5k
- Merge medio
- 1g 7h
- PR unite (30g)
- 340
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di mudler/LocalAI
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11995 · 1 commento ·
I maintainer di solito rispondono entro 3 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
mudler/LocalAI#11991 · 1 commento ·
I maintainer di solito rispondono entro 3 giorni
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasApertaenhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11348 · 1 commento ·
I maintainer di solito rispondono entro 3 giorni
-
feat: add automatic MCP transport selection for 2024-11-05 / 2025-03-26 / 2025-06-18 vs 2025-11-25Apertaenhancement
Difficoltà 3/5 1-2 giorni Idoneità per principianti 65/100
mudler/LocalAI#12262 · 2 commenti ·
I maintainer di solito rispondono entro 3 giorni
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
I maintainer di solito rispondono entro 3 giorni
Tutte le issue di mudler/LocalAI
Issue simili
-
priority: low 🌱 type: enhancement 💅🏼
Difficoltà 2/5 Mezza giornata Idoneità per principianti 84/100
nebari-dev/llm-serving-pack#199 ·
I maintainer di solito rispondono entro 3 giorni
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
-
area/helm kind/bug priority/backlog triage/accepted
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
lexfrei/cloudflare-tunnel-gateway-controller#889 ·
I maintainer di solito rispondono entro 1 giorno
-
bug difficulty: beginner documentation good first issue help wanted localization
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
wavefnd/wave-platform#140 ·
-
compiler/runtime
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
golang/go#81797 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno