Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Tool call written inside `<think>` (no `</think>`) is returned as `reasoning_content`: no `tool_calls`, `finish_reason=stop`

Chiusa
#804 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
45/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python

Direzione di ricerca

Start with OutputParser.feed() in serve/frontend.py and reproduce the issue using the minimal example in the report. Read the proposed tests in serve/test_reasoning_toolcall.py, then check the chat-template history cases in serve/chat_template.jinja and the pack-loading note. Done means real implicit tool calls are parsed consistently across streaming and non-streaming, quoted examples remain reasoning, and the parser, server, metrics, and template tests pass.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Summary

Qwen3-family models (seen with Qwen3.8-Flash-Next, GSQ-RCO IQ3_S) sometimes emit a complete, well-formed <tool_call> inside the thinking block without closing </think> first. serve/frontend.py → OutputParser only leaves the "reasoning" state on </think>, so the whole call ends up in reasoning_content. The client (an agent harness) receives tool_calls: [], empty content and finish_reason: "stop", and the agent loop silently stops: the call is never executed.

  • Frequency observed: 4 of ~1050 responses in real agent use (Hermes Agent, OpenAI Chat Completions API, streaming).
  • Engine 0.1.38, serve/server.py, thinking mode on, official Qwen thinking sampling (temperature 1.0, top_p 0.95, top_k 20, presence 0). Not a sampling issue.

What the model generated (shape of all 4 cases)

...the reasoning text, then the model decides to act.

<tool_call>
<function=execute_code>
<parameter=code>
print("hello")
</parameter>
</function>
</tool_call><|im_end|>

No </think> anywhere. The call itself is complete and parseable.

Minimal reproduction (no GPU)

from serve.frontend import OutputParser

text = ("I will run it now.\n\n<tool_call>\n<function=execute_code>\n"
        "<parameter=code>\nprint(1)\n</parameter>\n</function>\n</tool_call>")
p = OutputParser(thinking=True, tools=None, stream_tools=True)
evs = p.feed(text) + p.finish()
print([e.kind for e in evs])
# current main: ['reasoning']: no 'tool_call' event

Root cause

In OutputParser.feed(), the "reasoning" branch searches only for THINK_END (</think>). CALL_START is only looked for in the "content" state. A call generated before </think> is never seen as a call.

Prior art: vLLM

  • vLLM PR #35687 ("[Bugfix] Treat <tool_call> as implicit reasoning end in Qwen3 parser", merged 2026-04-24) fixed the same bug. In current vLLM (vllm/parser/qwen3.py, the grammar in qwen3_config()), the transition (REASONING, "TOOL_START") -> TOOL_PREAMBLE emits REASONING_END + TOOL_CALL_START, so any <tool_call> ends the reasoning, unconditionally.
  • vLLM PR #59821 ("Preserve quoted Qwen3 tool markup in reasoning") is still open. It addresses the false-positive side: a model that only mentions <tool_call> in its reasoning (quoting the format, explaining it) should not trigger a call.
  • vLLM's rule that ignores paired <tool_call> tokens applies to prompt tokens, not generated output, so it is not relevant here.

Proposed fix (implemented and tested locally, ready as a PR)

Adopt vLLM's idea (an implicit end of reasoning), but with a stricter trigger so quoted markup stays reasoning. In the "reasoning" state, a generated <tool_call> ends the thinking only when all three hold:

  1. it is at the start of a line (right after \n, or the very first text of the thinking);
  2. it is outside a fenced code block (``` or ~~~) of the reasoning;
  3. it is followed, after only whitespace, by <function= (the start of a real call body).

The text before it stays reasoning_content, and the call goes through the normal "call" state, so it is extracted in streaming (stream_tools) and non-streaming paths alike. While the text after <tool_call> is still only whitespace or a prefix of <function=, it is held back (the same approach as for partial tags), so streamed and whole outputs are identical.

Why stricter than vLLM: an unconditional rule turns things like "I must emit `<tool_call>` next", "The format is <tool_call>…" or an indented or >-quoted example into a bogus call. All real failures had the exact shape \n\n<tool_call>\n<function=NAME>, so the three conditions cost nothing in recall.

Edge case: if an implicit call never completes (output truncated), it is returned as reasoning, as before, not as content.

Observability: each implicit end is logged ([strata] implicit end of thinking: N tool call(s) written inside the thinking without </think> were read as calls) and counted in GET /metrics (totals.implicit_reasoning_ends plus a per-request field), so the rate can be tracked per model and quant.

Defense in depth: chat template

In serve/chat_template.jinja, when an earlier assistant message has a call left inside the thinking, </think> is now closed before the call when the history is rendered. This applies when:

  • content starts with an open <think> and contains a line-start <tool_call>\n<function=, or
  • there are no tool_calls and reasoning_content ends with such a call.

This way the model is never shown a call inside <think> as an example to imitate. All 10 cases in chat_golden.json render byte-identical. (Reference: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates.)

Note: the server loads the template from the pack (<pack>/tokenizer/chat_template.jinja), not from serve/. A template fix must also reach the pack, or strata_pack must copy it.

Tests

New serve/test_reasoning_toolcall.py (12 tests):

  • The 4 real failures (sanitized fixture): each yields exactly 1 tool_call, with the reasoning preserved. Checked when fed 1 char at a time, in random chunk sizes, and whole, with stream_tools on and off. All 4 yield 0 calls on current main.

  • 10 synthetic mention-only controls, none of which yield a call:

    • backticks mid-sentence;
    • a whole call written mid-sentence;
    • backticks at line start;
    • a line-start tag with no <function=;
    • another tag after it;
    • indented;
    • a markdown > quote;
    • inside ``` and ~~~ code blocks;
    • a bare tag before </think>.

    They also stay correct when followed by </think> and a real call.

  • Server-level: finish_reason=tool_calls, the log line, and the /metrics counter.

  • Template: chat_golden.json matches 10/10, and both history repairs are covered.

  • Full serve/test_*.py suite: 209 pass, 9 skipped (environment: Windows-only, no jsonschema, no pack tokenizer).

Side finding: agent harnesses and preserve_thinking

Hermes Agent strips reasoning_content from the history for generic OpenAI-compatible endpoints. It keeps it only for DeepSeek, Kimi and MiMo, or with model.reasoning_echo: true in its config (source: agent/message_sanitization.py). /metrics agrees: after a 9,977-token reply, the next prompt grew by only 64 tokens. With the default preserve_thinking=true, past assistant turns therefore render as empty <think>\n\n</think> blocks, and prefix-cache reuse stops at the last assistant turn. This is not a bug in Strata, but worth documenting for users of agent harnesses.

Keywords (for search)

tool_call inside think, missing </think>, reasoning_content contains tool_call, empty tool_calls, finish_reason stop instead of tool_calls, agent stops after reasoning, Qwen3 implicit reasoning end, OutputParser, Qwen3.8-Flash-Next, IQ3_S, Hermes Agent, vLLM #35687.

Image
Lingua principale
C++
Stelle
11.6k
Fork
1k
Merge medio
7h 46m
PR unite (30g)
30

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di Niko1221/Strata

Tutte le issue di Niko1221/Strata

Issue simili

Altre issue su C++

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.