Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Tool call arguments are empty on /responses but correct on /chat/completions (same model, same config)

Đã đóng
#11,635 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
55/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
go
Lĩnh vực
api, backend

Hướng nghiên cứu

Bắt đầu tại điểm vào /v1/responses và so sánh cách lắp ráp function_call của nó với /v1/chat/completions, bằng cách tái hiện các yêu cầu curl được cung cấp khi đã bật ghi log DEBUG. Được xem là hoàn tất khi các mục function_call của Responses API giữ nguyên các đối số do backend tạo ra, bao gồm cả với các call IDs có tiền tố fc_-, và phần kiểm thử hồi quy liên quan đều đạt.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

area/api bug unconfirmed

Tool call arguments are empty on /responses but correct on /chat/completions (same model, same config)

LocalAI version:

LocalAI v4.8.2

Environment, CPU architecture, OS, and Version:

  • Docker on Ubuntu, x86_64
  • GPU: NVIDIA GeForce RTX 5090
  • Backend: llama-cpp
  • Not a VM (bare metal)

Describe the bug

With the same model and the same model config, tool calls returned through /v1/chat/completions contain complete arguments, while tool calls returned through /v1/responses contain an empty arguments object ("{}").

The tool name is present and correct on both paths — only the arguments are lost. Because the name survives, the client receives a structurally valid tool call that then fails its own schema validation (The required parameter 'command' is missing), and the agent retries the same call repeatedly until the turn is abandoned.

To Reproduce

Model config (/models/qwen3.8-27b-q4.yaml):

backend: llama-cpp
context_size: 204800
flash_attention: true
cache_type_k: "q8_0"
cache_type_v: "q8_0"
function:
    automatic_tool_parsing_fallback: false
    grammar:
        disable: false
known_usecases:
    - chat
    - vision
mmproj: llama-cpp/mmproj/qwen3.8-27b/mmproj-Qwen3.8-27B-Q8_0.gguf
name: qwen3.8-27b-q4
options:
    - use_jinja:true
    - "--chat-template-file:/models/chat_template.jinja"
parameters:
    min_p: 0
    model: llama-cpp/models/qwen3.8-27b/Qwen3.8-27B-Q4_K_M.gguf
    presence_penalty: 0
    repeat_penalty: 1
    temperature: 1
    top_k: 20
    top_p: 0.95
chat_template_kwargs:
    tool_call_format: "json"
template:
    use_tokenizer_template: false

chat_template.jinja is a community-maintained Jinja template for Qwen 3.x (froggeric/Qwen-Fixed-Chat-Templates, v22.2), loaded via the --chat-template-file passthrough.

Step 1 — /v1/chat/completions (correct):

curl -s http://<localai-host>/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b-q4",
    "messages": [{"role": "user", "content": "run ls -la"}],
    "tools": [{
      "type": "function",
      "function": {
        "name": "Bash",
        "parameters": {
          "type": "object",
          "properties": {"command": {"type": "string"}},
          "required": ["command"]
        }
      }
    }],
    "stream": false
  }'

Arguments are complete:

{
  "choices": [{
    "finish_reason": "tool_calls",
    "message": {
      "content": "",
      "role": "assistant",
      "tool_calls": [{
        "index": 0,
        "id": "N3hviczmpFmHuxC9d3tewiFzOJo8mA9z",
        "type": "function",
        "function": {
          "name": "Bash",
          "arguments": "{\"command\":\"ls -la\"}"
        }
      }],
      "reasoning_content": "User wants to run ls -la.\n"
    }
  }]
}

Step 2 — /v1/responses (arguments empty):

curl -s http://<localai-host>/v1/responses \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.8-27b-q4",
    "input": [{
      "type": "message",
      "role": "user",
      "content": [{"type": "input_text", "text": "run ls -la"}]
    }],
    "tools": [{
      "type": "function",
      "name": "Bash",
      "parameters": {
        "type": "object",
        "properties": {"command": {"type": "string"}},
        "required": ["command"]
      }
    }],
    "stream": false
  }'

The function_call item comes back with "arguments": "{}".

Observed across a real agent session, two shapes appear in the same conversation:

{"type": "function_call",
 "call_id": "KoYQ6Mgq3ud4GFH4A5TXzzi3xQ0L6Z4W",
 "name": "Glob",
 "arguments": "{\"pattern\": \"src/app/components/case-form/**\"}"}

{"type": "function_call",
 "call_id": "fc_0c37f48a-ae9b-4bad-867e-d3f5b0772086",
 "name": "Glob",
 "arguments": "{}"}

Every call whose call_id is a plain random string carries correct arguments. Every call whose call_id has an fc_ prefix has empty arguments. I have not traced where the fc_ identifiers originate, so this is reported as a consistent correlation, not a diagnosis.

In one turn, five consecutive Read calls were emitted with empty arguments and fc_-prefixed ids before the turn was abandoned.

Expected behavior

/v1/responses should return function_call items whose arguments match what the backend produced, identically to the tool_calls returned by /v1/chat/completions for an equivalent request.

Logs

Backend logging is enabled on this instance. I can also attach backend traces for a failing /v1/responses request if that helps isolate whether the arguments are lost in the backend response or in the Responses API adapter.

Additional context

This may share a cause with #9334 ("Gemma 4 Tool Response is not returned as expected", v4.1.3), which reports tool call responses visible in backend traces but absent from the API response, along with the same five retries. That report was against /v1/chat/completions, whereas here /v1/chat/completions is correct and /v1/responses is not — so they may be separate code paths with a shared underlying problem.

Why this matters in practice: gateways that translate Anthropic's /v1/messages into an OpenAI-shaped request (LiteLLM, Bifrost) route to /v1/responses whenever the target model has reasoning enabled. Qwen 3.8 always reasons and cannot have thinking disabled, so every request from an Anthropic-format client such as Claude Code lands on the affected path. Opting out at the gateway layer is also unreliable — see BerriAI/litellm#23841, which documents three separate places where the opt-out setting is ignored. The practical effect is that tool use through Anthropic-format clients does not work against LocalAI unless the gateway is forced onto /chat/completions by other means.

Ngôn ngữ chính
Go
Star
49.2k
Fork
4.5k
Merge trung bình
1 ngày 7 giờ
Pull request đã merge (30 ngày)
357

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của mudler/LocalAI

Tất cả issue của mudler/LocalAI

Issue tương tự

Thêm issue về Go

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.