Tool call arguments are empty on /responses but correct on /chat/completions (same model, same config)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 55/100
Direzione di ricerca
Inizia dal punto di ingresso /v1/responses e confronta l'assemblaggio di function_call con /v1/chat/completions, riproducendo le richieste curl fornite con il logging DEBUG abilitato. Il lavoro è completato quando gli elementi function_call della Responses API conservano gli argomenti prodotti dal backend, anche per i call IDs con prefisso fc_-, e la copertura di regressione pertinente ha esito positivo.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Tool call arguments are empty on /responses but correct on /chat/completions (same model, same config)
LocalAI version:
LocalAI v4.8.2
Environment, CPU architecture, OS, and Version:
- Docker on Ubuntu, x86_64
- GPU: NVIDIA GeForce RTX 5090
- Backend:
llama-cpp - Not a VM (bare metal)
Describe the bug
With the same model and the same model config, tool calls returned through /v1/chat/completions contain complete arguments, while tool calls returned through /v1/responses contain an empty arguments object ("{}").
The tool name is present and correct on both paths — only the arguments are lost. Because the name survives, the client receives a structurally valid tool call that then fails its own schema validation (The required parameter 'command' is missing), and the agent retries the same call repeatedly until the turn is abandoned.
To Reproduce
Model config (/models/qwen3.8-27b-q4.yaml):
backend: llama-cpp
context_size: 204800
flash_attention: true
cache_type_k: "q8_0"
cache_type_v: "q8_0"
function:
automatic_tool_parsing_fallback: false
grammar:
disable: false
known_usecases:
- chat
- vision
mmproj: llama-cpp/mmproj/qwen3.8-27b/mmproj-Qwen3.8-27B-Q8_0.gguf
name: qwen3.8-27b-q4
options:
- use_jinja:true
- "--chat-template-file:/models/chat_template.jinja"
parameters:
min_p: 0
model: llama-cpp/models/qwen3.8-27b/Qwen3.8-27B-Q4_K_M.gguf
presence_penalty: 0
repeat_penalty: 1
temperature: 1
top_k: 20
top_p: 0.95
chat_template_kwargs:
tool_call_format: "json"
template:
use_tokenizer_template: false
chat_template.jinja is a community-maintained Jinja template for Qwen 3.x (froggeric/Qwen-Fixed-Chat-Templates, v22.2), loaded via the --chat-template-file passthrough.
Step 1 — /v1/chat/completions (correct):
curl -s http://<localai-host>/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b-q4",
"messages": [{"role": "user", "content": "run ls -la"}],
"tools": [{
"type": "function",
"function": {
"name": "Bash",
"parameters": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"]
}
}
}],
"stream": false
}'
Arguments are complete:
{
"choices": [{
"finish_reason": "tool_calls",
"message": {
"content": "",
"role": "assistant",
"tool_calls": [{
"index": 0,
"id": "N3hviczmpFmHuxC9d3tewiFzOJo8mA9z",
"type": "function",
"function": {
"name": "Bash",
"arguments": "{\"command\":\"ls -la\"}"
}
}],
"reasoning_content": "User wants to run ls -la.\n"
}
}]
}
Step 2 — /v1/responses (arguments empty):
curl -s http://<localai-host>/v1/responses \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b-q4",
"input": [{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "run ls -la"}]
}],
"tools": [{
"type": "function",
"name": "Bash",
"parameters": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"]
}
}],
"stream": false
}'
The function_call item comes back with "arguments": "{}".
Observed across a real agent session, two shapes appear in the same conversation:
{"type": "function_call",
"call_id": "KoYQ6Mgq3ud4GFH4A5TXzzi3xQ0L6Z4W",
"name": "Glob",
"arguments": "{\"pattern\": \"src/app/components/case-form/**\"}"}
{"type": "function_call",
"call_id": "fc_0c37f48a-ae9b-4bad-867e-d3f5b0772086",
"name": "Glob",
"arguments": "{}"}
Every call whose call_id is a plain random string carries correct arguments. Every call whose call_id has an fc_ prefix has empty arguments. I have not traced where the fc_ identifiers originate, so this is reported as a consistent correlation, not a diagnosis.
In one turn, five consecutive Read calls were emitted with empty arguments and fc_-prefixed ids before the turn was abandoned.
Expected behavior
/v1/responses should return function_call items whose arguments match what the backend produced, identically to the tool_calls returned by /v1/chat/completions for an equivalent request.
Logs
Backend logging is enabled on this instance. I can also attach backend traces for a failing /v1/responses request if that helps isolate whether the arguments are lost in the backend response or in the Responses API adapter.
Additional context
This may share a cause with #9334 ("Gemma 4 Tool Response is not returned as expected", v4.1.3), which reports tool call responses visible in backend traces but absent from the API response, along with the same five retries. That report was against /v1/chat/completions, whereas here /v1/chat/completions is correct and /v1/responses is not — so they may be separate code paths with a shared underlying problem.
Why this matters in practice: gateways that translate Anthropic's /v1/messages into an OpenAI-shaped request (LiteLLM, Bifrost) route to /v1/responses whenever the target model has reasoning enabled. Qwen 3.8 always reasons and cannot have thinking disabled, so every request from an Anthropic-format client such as Claude Code lands on the affected path. Opting out at the gateway layer is also unreliable — see BerriAI/litellm#23841, which documents three separate places where the opt-out setting is ignored. The practical effect is that tool use through Anthropic-format clients does not work against LocalAI unless the gateway is forced onto /chat/completions by other means.
- Lingua principale
- Go
- Stelle
- 49.2k
- Fork
- 4.5k
- Merge medio
- 1g 7h
- PR unite (30g)
- 362
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di mudler/LocalAI
-
bug unconfirmed
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
mudler/LocalAI#12337 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11995 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
mudler/LocalAI#11991 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasApertaenhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
mudler/LocalAI#11348 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
barholeurApertabug unconfirmed
Difficoltà 5/5 Più di una settimana Idoneità per principianti 10/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di mudler/LocalAI
Issue simili
-
area: global bug dx priority: low
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
grafana/mcp-grafana#1267 ·
I maintainer di solito rispondono entro 1 giorno
-
automation models
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
coverage-gap good-first-pattern help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
GoogleCloudPlatform/k8s-aibom#114 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
txn2/mcp-data-platform#1984 ·
I maintainer di solito rispondono entro 1 giorno