Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Tool call written inside `<think>` (no `</think>`) is returned as `reasoning_content`: no `tool_calls`, `finish_reason=stop`

Open
#804 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
45/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

Start with OutputParser.feed() in serve/frontend.py and reproduce the issue using the minimal example in the report. Read the proposed tests in serve/test_reasoning_toolcall.py, then check the chat-template history cases in serve/chat_template.jinja and the pack-loading note. Done means real implicit tool calls are parsed consistently across streaming and non-streaming, quoted examples remain reasoning, and the parser, server, metrics, and template tests pass.

Written by the indexing model from the issue text.

Description

Summary

Qwen3-family models (seen with Qwen3.8-Flash-Next, GSQ-RCO IQ3_S) sometimes emit a complete, well-formed <tool_call> inside the thinking block without closing </think> first. serve/frontend.py → OutputParser only leaves the "reasoning" state on </think>, so the whole call ends up in reasoning_content. The client (an agent harness) receives tool_calls: [], empty content and finish_reason: "stop", and the agent loop silently stops: the call is never executed.

  • Frequency observed: 4 of ~1050 responses in real agent use (Hermes Agent, OpenAI Chat Completions API, streaming).
  • Engine 0.1.38, serve/server.py, thinking mode on, official Qwen thinking sampling (temperature 1.0, top_p 0.95, top_k 20, presence 0). Not a sampling issue.

What the model generated (shape of all 4 cases)

...the reasoning text, then the model decides to act.

<tool_call>
<function=execute_code>
<parameter=code>
print("hello")
</parameter>
</function>
</tool_call><|im_end|>

No </think> anywhere. The call itself is complete and parseable.

Minimal reproduction (no GPU)

from serve.frontend import OutputParser

text = ("I will run it now.\n\n<tool_call>\n<function=execute_code>\n"
        "<parameter=code>\nprint(1)\n</parameter>\n</function>\n</tool_call>")
p = OutputParser(thinking=True, tools=None, stream_tools=True)
evs = p.feed(text) + p.finish()
print([e.kind for e in evs])
# current main: ['reasoning']: no 'tool_call' event

Root cause

In OutputParser.feed(), the "reasoning" branch searches only for THINK_END (</think>). CALL_START is only looked for in the "content" state. A call generated before </think> is never seen as a call.

Prior art: vLLM

  • vLLM PR #35687 ("[Bugfix] Treat <tool_call> as implicit reasoning end in Qwen3 parser", merged 2026-04-24) fixed the same bug. In current vLLM (vllm/parser/qwen3.py, the grammar in qwen3_config()), the transition (REASONING, "TOOL_START") -> TOOL_PREAMBLE emits REASONING_END + TOOL_CALL_START, so any <tool_call> ends the reasoning, unconditionally.
  • vLLM PR #59821 ("Preserve quoted Qwen3 tool markup in reasoning") is still open. It addresses the false-positive side: a model that only mentions <tool_call> in its reasoning (quoting the format, explaining it) should not trigger a call.
  • vLLM's rule that ignores paired <tool_call> tokens applies to prompt tokens, not generated output, so it is not relevant here.

Proposed fix (implemented and tested locally, ready as a PR)

Adopt vLLM's idea (an implicit end of reasoning), but with a stricter trigger so quoted markup stays reasoning. In the "reasoning" state, a generated <tool_call> ends the thinking only when all three hold:

  1. it is at the start of a line (right after \n, or the very first text of the thinking);
  2. it is outside a fenced code block (``` or ~~~) of the reasoning;
  3. it is followed, after only whitespace, by <function= (the start of a real call body).

The text before it stays reasoning_content, and the call goes through the normal "call" state, so it is extracted in streaming (stream_tools) and non-streaming paths alike. While the text after <tool_call> is still only whitespace or a prefix of <function=, it is held back (the same approach as for partial tags), so streamed and whole outputs are identical.

Why stricter than vLLM: an unconditional rule turns things like "I must emit `<tool_call>` next", "The format is <tool_call>…" or an indented or >-quoted example into a bogus call. All real failures had the exact shape \n\n<tool_call>\n<function=NAME>, so the three conditions cost nothing in recall.

Edge case: if an implicit call never completes (output truncated), it is returned as reasoning, as before, not as content.

Observability: each implicit end is logged ([strata] implicit end of thinking: N tool call(s) written inside the thinking without </think> were read as calls) and counted in GET /metrics (totals.implicit_reasoning_ends plus a per-request field), so the rate can be tracked per model and quant.

Defense in depth: chat template

In serve/chat_template.jinja, when an earlier assistant message has a call left inside the thinking, </think> is now closed before the call when the history is rendered. This applies when:

  • content starts with an open <think> and contains a line-start <tool_call>\n<function=, or
  • there are no tool_calls and reasoning_content ends with such a call.

This way the model is never shown a call inside <think> as an example to imitate. All 10 cases in chat_golden.json render byte-identical. (Reference: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates.)

Note: the server loads the template from the pack (<pack>/tokenizer/chat_template.jinja), not from serve/. A template fix must also reach the pack, or strata_pack must copy it.

Tests

New serve/test_reasoning_toolcall.py (12 tests):

  • The 4 real failures (sanitized fixture): each yields exactly 1 tool_call, with the reasoning preserved. Checked when fed 1 char at a time, in random chunk sizes, and whole, with stream_tools on and off. All 4 yield 0 calls on current main.

  • 10 synthetic mention-only controls, none of which yield a call:

    • backticks mid-sentence;
    • a whole call written mid-sentence;
    • backticks at line start;
    • a line-start tag with no <function=;
    • another tag after it;
    • indented;
    • a markdown > quote;
    • inside ``` and ~~~ code blocks;
    • a bare tag before </think>.

    They also stay correct when followed by </think> and a real call.

  • Server-level: finish_reason=tool_calls, the log line, and the /metrics counter.

  • Template: chat_golden.json matches 10/10, and both history repairs are covered.

  • Full serve/test_*.py suite: 209 pass, 9 skipped (environment: Windows-only, no jsonschema, no pack tokenizer).

Side finding: agent harnesses and preserve_thinking

Hermes Agent strips reasoning_content from the history for generic OpenAI-compatible endpoints. It keeps it only for DeepSeek, Kimi and MiMo, or with model.reasoning_echo: true in its config (source: agent/message_sanitization.py). /metrics agrees: after a 9,977-token reply, the next prompt grew by only 64 tokens. With the default preserve_thinking=true, past assistant turns therefore render as empty <think>\n\n</think> blocks, and prefix-cache reuse stops at the last assistant turn. This is not a bug in Strata, but worth documenting for users of agent harnesses.

Keywords (for search)

tool_call inside think, missing </think>, reasoning_content contains tool_call, empty tool_calls, finish_reason stop instead of tool_calls, agent stops after reasoning, Qwen3 implicit reasoning end, OutputParser, Qwen3.8-Flash-Next, IQ3_S, Hermes Agent, vLLM #35687.

Image
Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.