Tool call arguments are empty on /responses but correct on /chat/completions (same model, same config)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 55/100
Línea de trabajo
Comienza en el punto de entrada /v1/responses y compara el ensamblado de function_call con /v1/chat/completions, reproduciendo las solicitudes curl proporcionadas con el registro DEBUG habilitado. Se considera terminado cuando los elementos function_call de la Responses API conservan los argumentos generados por el backend, incluidos los call IDs con prefijo fc_-, y pasa la cobertura de regresión pertinente.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Tool call arguments are empty on /responses but correct on /chat/completions (same model, same config)
LocalAI version:
LocalAI v4.8.2
Environment, CPU architecture, OS, and Version:
- Docker on Ubuntu, x86_64
- GPU: NVIDIA GeForce RTX 5090
- Backend:
llama-cpp - Not a VM (bare metal)
Describe the bug
With the same model and the same model config, tool calls returned through /v1/chat/completions contain complete arguments, while tool calls returned through /v1/responses contain an empty arguments object ("{}").
The tool name is present and correct on both paths — only the arguments are lost. Because the name survives, the client receives a structurally valid tool call that then fails its own schema validation (The required parameter 'command' is missing), and the agent retries the same call repeatedly until the turn is abandoned.
To Reproduce
Model config (/models/qwen3.8-27b-q4.yaml):
backend: llama-cpp
context_size: 204800
flash_attention: true
cache_type_k: "q8_0"
cache_type_v: "q8_0"
function:
automatic_tool_parsing_fallback: false
grammar:
disable: false
known_usecases:
- chat
- vision
mmproj: llama-cpp/mmproj/qwen3.8-27b/mmproj-Qwen3.8-27B-Q8_0.gguf
name: qwen3.8-27b-q4
options:
- use_jinja:true
- "--chat-template-file:/models/chat_template.jinja"
parameters:
min_p: 0
model: llama-cpp/models/qwen3.8-27b/Qwen3.8-27B-Q4_K_M.gguf
presence_penalty: 0
repeat_penalty: 1
temperature: 1
top_k: 20
top_p: 0.95
chat_template_kwargs:
tool_call_format: "json"
template:
use_tokenizer_template: false
chat_template.jinja is a community-maintained Jinja template for Qwen 3.x (froggeric/Qwen-Fixed-Chat-Templates, v22.2), loaded via the --chat-template-file passthrough.
Step 1 — /v1/chat/completions (correct):
curl -s http://<localai-host>/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b-q4",
"messages": [{"role": "user", "content": "run ls -la"}],
"tools": [{
"type": "function",
"function": {
"name": "Bash",
"parameters": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"]
}
}
}],
"stream": false
}'
Arguments are complete:
{
"choices": [{
"finish_reason": "tool_calls",
"message": {
"content": "",
"role": "assistant",
"tool_calls": [{
"index": 0,
"id": "N3hviczmpFmHuxC9d3tewiFzOJo8mA9z",
"type": "function",
"function": {
"name": "Bash",
"arguments": "{\"command\":\"ls -la\"}"
}
}],
"reasoning_content": "User wants to run ls -la.\n"
}
}]
}
Step 2 — /v1/responses (arguments empty):
curl -s http://<localai-host>/v1/responses \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b-q4",
"input": [{
"type": "message",
"role": "user",
"content": [{"type": "input_text", "text": "run ls -la"}]
}],
"tools": [{
"type": "function",
"name": "Bash",
"parameters": {
"type": "object",
"properties": {"command": {"type": "string"}},
"required": ["command"]
}
}],
"stream": false
}'
The function_call item comes back with "arguments": "{}".
Observed across a real agent session, two shapes appear in the same conversation:
{"type": "function_call",
"call_id": "KoYQ6Mgq3ud4GFH4A5TXzzi3xQ0L6Z4W",
"name": "Glob",
"arguments": "{\"pattern\": \"src/app/components/case-form/**\"}"}
{"type": "function_call",
"call_id": "fc_0c37f48a-ae9b-4bad-867e-d3f5b0772086",
"name": "Glob",
"arguments": "{}"}
Every call whose call_id is a plain random string carries correct arguments. Every call whose call_id has an fc_ prefix has empty arguments. I have not traced where the fc_ identifiers originate, so this is reported as a consistent correlation, not a diagnosis.
In one turn, five consecutive Read calls were emitted with empty arguments and fc_-prefixed ids before the turn was abandoned.
Expected behavior
/v1/responses should return function_call items whose arguments match what the backend produced, identically to the tool_calls returned by /v1/chat/completions for an equivalent request.
Logs
Backend logging is enabled on this instance. I can also attach backend traces for a failing /v1/responses request if that helps isolate whether the arguments are lost in the backend response or in the Responses API adapter.
Additional context
This may share a cause with #9334 ("Gemma 4 Tool Response is not returned as expected", v4.1.3), which reports tool call responses visible in backend traces but absent from the API response, along with the same five retries. That report was against /v1/chat/completions, whereas here /v1/chat/completions is correct and /v1/responses is not — so they may be separate code paths with a shared underlying problem.
Why this matters in practice: gateways that translate Anthropic's /v1/messages into an OpenAI-shaped request (LiteLLM, Bifrost) route to /v1/responses whenever the target model has reasoning enabled. Qwen 3.8 always reasons and cannot have thinking disabled, so every request from an Anthropic-format client such as Claude Code lands on the affected path. Opting out at the gateway layer is also unreliable — see BerriAI/litellm#23841, which documents three separate places where the opt-out setting is ignored. The practical effect is that tool use through Anthropic-format clients does not work against LocalAI unless the gateway is forced onto /chat/completions by other means.
- Lenguaje dominante
- Go
- Estrellas
- 49.2k
- Forks
- 4.5k
- Merge medio
- 1 d 7 h
- PR fusionados (30 d)
- 362
Preparar el entorno
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de mudler/LocalAI
-
bug unconfirmed
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
mudler/LocalAI#12337 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
mudler/LocalAI#11995 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
mudler/LocalAI#11991 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
fish-speech: make compile:true usable on Blackwell sm_121 by honouring the CUDA toolkit's ptxasAbiertoenhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
mudler/LocalAI#11348 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
barholeurAbiertobug unconfirmed
Dificultad 5/5 Más de una semana Aptitud para principiantes 10/100
Los mantenedores suelen responder en 1 día
Todos los issues de mudler/LocalAI
Issues similares
-
agent-butler-finding chore
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
jordansmall/spindrift#4146 ·
Los mantenedores suelen responder en 1 día
-
security
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
IBM/ibmcloud-volume-file-vpc#119 ·
-
security
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100
IBM/networking-go-sdk#339 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
kubernetes-sigs/mcp-lifecycle-operator#439 ·
Los mantenedores suelen responder en 1 día
-
area: global bug dx priority: low
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día