Tool call written inside `<think>` (no `</think>`) is returned as `reasoning_content`: no `tool_calls`, `finish_reason=stop`
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 45/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- backend-api-design
Research direction
Start with OutputParser.feed() in serve/frontend.py and reproduce the issue using the minimal example in the report. Read the proposed tests in serve/test_reasoning_toolcall.py, then check the chat-template history cases in serve/chat_template.jinja and the pack-loading note. Done means real implicit tool calls are parsed consistently across streaming and non-streaming, quoted examples remain reasoning, and the parser, server, metrics, and template tests pass.
Written by the indexing model from the issue text.
Description
Summary
Qwen3-family models (seen with Qwen3.8-Flash-Next, GSQ-RCO IQ3_S) sometimes emit a complete, well-formed <tool_call> inside the thinking block without closing </think> first. serve/frontend.py → OutputParser only leaves the "reasoning" state on </think>, so the whole call ends up in reasoning_content. The client (an agent harness) receives tool_calls: [], empty content and finish_reason: "stop", and the agent loop silently stops: the call is never executed.
- Frequency observed: 4 of ~1050 responses in real agent use (Hermes Agent, OpenAI Chat Completions API, streaming).
- Engine 0.1.38,
serve/server.py, thinking mode on, official Qwen thinking sampling (temperature 1.0, top_p 0.95, top_k 20, presence 0). Not a sampling issue.
What the model generated (shape of all 4 cases)
...the reasoning text, then the model decides to act.
<tool_call>
<function=execute_code>
<parameter=code>
print("hello")
</parameter>
</function>
</tool_call><|im_end|>
No </think> anywhere. The call itself is complete and parseable.
Minimal reproduction (no GPU)
from serve.frontend import OutputParser
text = ("I will run it now.\n\n<tool_call>\n<function=execute_code>\n"
"<parameter=code>\nprint(1)\n</parameter>\n</function>\n</tool_call>")
p = OutputParser(thinking=True, tools=None, stream_tools=True)
evs = p.feed(text) + p.finish()
print([e.kind for e in evs])
# current main: ['reasoning']: no 'tool_call' event
Root cause
In OutputParser.feed(), the "reasoning" branch searches only for THINK_END (</think>). CALL_START is only looked for in the "content" state. A call generated before </think> is never seen as a call.
Prior art: vLLM
- vLLM PR #35687 ("[Bugfix] Treat
<tool_call>as implicit reasoning end in Qwen3 parser", merged 2026-04-24) fixed the same bug. In current vLLM (vllm/parser/qwen3.py, the grammar inqwen3_config()), the transition(REASONING, "TOOL_START") -> TOOL_PREAMBLEemitsREASONING_END+TOOL_CALL_START, so any<tool_call>ends the reasoning, unconditionally. - vLLM PR #59821 ("Preserve quoted Qwen3 tool markup in reasoning") is still open. It addresses the false-positive side: a model that only mentions
<tool_call>in its reasoning (quoting the format, explaining it) should not trigger a call. - vLLM's rule that ignores paired
<tool_call>tokens applies to prompt tokens, not generated output, so it is not relevant here.
Proposed fix (implemented and tested locally, ready as a PR)
Adopt vLLM's idea (an implicit end of reasoning), but with a stricter trigger so quoted markup stays reasoning. In the "reasoning" state, a generated <tool_call> ends the thinking only when all three hold:
- it is at the start of a line (right after
\n, or the very first text of the thinking); - it is outside a fenced code block (
```or~~~) of the reasoning; - it is followed, after only whitespace, by
<function=(the start of a real call body).
The text before it stays reasoning_content, and the call goes through the normal "call" state, so it is extracted in streaming (stream_tools) and non-streaming paths alike. While the text after <tool_call> is still only whitespace or a prefix of <function=, it is held back (the same approach as for partial tags), so streamed and whole outputs are identical.
Why stricter than vLLM: an unconditional rule turns things like "I must emit `<tool_call>` next", "The format is <tool_call>…" or an indented or >-quoted example into a bogus call. All real failures had the exact shape \n\n<tool_call>\n<function=NAME>, so the three conditions cost nothing in recall.
Edge case: if an implicit call never completes (output truncated), it is returned as reasoning, as before, not as content.
Observability: each implicit end is logged ([strata] implicit end of thinking: N tool call(s) written inside the thinking without </think> were read as calls) and counted in GET /metrics (totals.implicit_reasoning_ends plus a per-request field), so the rate can be tracked per model and quant.
Defense in depth: chat template
In serve/chat_template.jinja, when an earlier assistant message has a call left inside the thinking, </think> is now closed before the call when the history is rendered. This applies when:
contentstarts with an open<think>and contains a line-start<tool_call>\n<function=, or- there are no
tool_callsandreasoning_contentends with such a call.
This way the model is never shown a call inside <think> as an example to imitate. All 10 cases in chat_golden.json render byte-identical. (Reference: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates.)
Note: the server loads the template from the pack (<pack>/tokenizer/chat_template.jinja), not from serve/. A template fix must also reach the pack, or strata_pack must copy it.
Tests
New serve/test_reasoning_toolcall.py (12 tests):
-
The 4 real failures (sanitized fixture): each yields exactly 1
tool_call, with the reasoning preserved. Checked when fed 1 char at a time, in random chunk sizes, and whole, withstream_toolson and off. All 4 yield 0 calls on currentmain. -
10 synthetic mention-only controls, none of which yield a call:
- backticks mid-sentence;
- a whole call written mid-sentence;
- backticks at line start;
- a line-start tag with no
<function=; - another tag after it;
- indented;
- a markdown
>quote; - inside
```and~~~code blocks; - a bare tag before
</think>.
They also stay correct when followed by
</think>and a real call. -
Server-level:
finish_reason=tool_calls, the log line, and the/metricscounter. -
Template:
chat_golden.jsonmatches 10/10, and both history repairs are covered. -
Full
serve/test_*.pysuite: 209 pass, 9 skipped (environment: Windows-only, no jsonschema, no pack tokenizer).
Side finding: agent harnesses and preserve_thinking
Hermes Agent strips reasoning_content from the history for generic OpenAI-compatible endpoints. It keeps it only for DeepSeek, Kimi and MiMo, or with model.reasoning_echo: true in its config (source: agent/message_sanitization.py). /metrics agrees: after a 9,977-token reply, the next prompt grew by only 64 tokens. With the default preserve_thinking=true, past assistant turns therefore render as empty <think>\n\n</think> blocks, and prefix-cache reuse stops at the last assistant turn. This is not a bug in Strata, but worth documenting for users of agent harnesses.
Keywords (for search)
tool_call inside think, missing </think>, reasoning_content contains tool_call, empty tool_calls, finish_reason stop instead of tool_calls, agent stops after reasoning, Qwen3 implicit reasoning end, OutputParser, Qwen3.8-Flash-Next, IQ3_S, Hermes Agent, vLLM #35687.
- Dominant language
- C++
- Stars
- 11.6k
- Forks
- 1k
- Avg merge
- 7h 46m
- Merged PRs (30d)
- 30
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Niko1221/Strata
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Niko1221/Strata#974 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 73/100
EchoTools/nevr-runtime#116 ·
Maintainers usually reply within 1 day
-
code-quality libc++
Difficulty 1/5 Under an hour Newbie friendliness 82/100
llvm/llvm-project#229284 ·
Maintainers usually reply within 1 day
-
test-issue
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
llvm/offload-test-suite#1557 ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
iOS: hidden scale bar invalidates its intrinsic content size on every layout pass of MLNMapViewOpen
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
maplibre/maplibre-native#4723 ·
Maintainers usually reply within 1 day