Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Qwen3.5 llama-cpp tool calling fails when function has "strict": true

オープン
#11,709 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 3 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
52/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
go

調査の方向性

/v1/chat/completions のエントリポイントから開始し、llama-cpp バックエンドが strict: true の場合とそうでない場合に function tools をどのように処理するかを追跡します。issue にある 2 つのリクエストをモデル qwen3.5-9b-glm5.1-distill-v1 で再現し、その後、strict リクエストが空のレスポンスとリトライループではなくツール呼び出しを返すことを確認します。

索引モデルが issue の本文から書いたものです。

説明

area/api area/llama.cpp bug confirmed

LocalAI version:

LocalAI version: v4.9.0

Commit:
f7ad3f70eb5d8a0ddf80e08557f0d7df28cf032e

Container image:
localai/localai:latest-gpu-nvidia-cuda-12

LocalAI is running as a Kubernetes deployment.


Environment, CPU architecture, OS, and Version:

LocalAI is running in Kubernetes.

Uname:

Linux k8scp 5.14.0-611.47.1.el9_7.x86_64 #1 SMP PREEMPT_DYNAMIC Wed Apr 8 12:18:23 UTC 2026 x86_64 x86_64 x86_64 GNU/Linux

Model:

qwen3.5-9b-glm5.1-distill-v1

Model file:

Qwen3.5-9B-GLM5.1-Distill-v1-Q4_K_M.gguf

Backend:

llama-cpp

Effective runtime settings reported by LocalAI:

context=2048
n_batch=512
n_gpu_layers=0
parallel="1"
flash_attention="auto"
f16=false

The model is therefore running CPU-only for this reproduction
(n_gpu_layers=0).

CPU is an Intel Core i7-3770 @ 3.40 GHz.

Client testing was performed using the OpenAI-compatible
/v1/chat/completions endpoint.

The issue was originally discovered using OpenAI Agents Python SDK
0.22.0, but it can be reproduced independently of the Agents SDK
using a normal OpenAI-compatible Chat Completions request.


Describe the bug

OpenAI-compatible function/tool calling works with the
qwen3.5-9b-glm5.1-distill-v1 model using the llama-cpp backend, but
adding:

"strict": true

to the function definition causes tool calling to fail.

Without strict: true, the model correctly returns:

finish_reason: tool_calls
tool_calls: transfer_to_customer_retention_agent

With strict: true, the same request instead returns:

finish_reason: stop
content: ''
tool_calls: None
reasoning: None

LocalAI returns HTTP 200 in both cases.

The LocalAI logs show repeated:

Backend returned empty response, retrying

for the failing request.

This was initially noticed because OpenAI Agents SDK handoffs are
represented as strict function tools, which prevents native Agents SDK
handoffs from working with this LocalAI/model combination.

The issue can, however, be reproduced without using the Agents SDK.


To Reproduce

Use the following model:

qwen3.5-9b-glm5.1-distill-v1

with the llama-cpp backend.

1. Tool without strict: true - works

Send:

curl http://<LOCALAI_HOST>:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "qwen3.5-9b-glm5.1-distill-v1",
    "messages": [
      {
        "role": "system",
        "content": "You are a customer service agent. If the customer wants to close or cancel their account, you must call transfer_to_customer_retention_agent."
      },
      {
        "role": "user",
        "content": "I want to cancel my order and account. You delayed my order for the 3rd time!"
      }
    ],
    "tools": [
      {
        "type": "function",
        "function": {
          "name": "transfer_to_customer_retention_agent",
          "description": "Transfer the conversation to the customer retention specialist.",
          "parameters": {
            "type": "object",
            "properties": {},
            "required": [],
            "additionalProperties": false
          }
        }
      }
    ],
    "tool_choice": "auto",
    "max_tokens": 1500,
    "reasoning_effort": "none"
  }'

This correctly results in an OpenAI-compatible tool call.

Using the OpenAI Python client, the relevant response is:

finish_reason: tool_calls
content: ''
tool_calls:
[
  ChatCompletionMessageFunctionToolCall(
    function=Function(
      arguments='{}',
      name='transfer_to_customer_retention_agent'
    ),
    type='function'
  )
]
2. Add only strict: true - fails

Repeat the same request but change the function definition to:

{
  "type": "function",
  "function": {
    "name": "transfer_to_customer_retention_agent",
    "description": "Transfer the conversation to the customer retention specialist.",
    "parameters": {
      "type": "object",
      "properties": {},
      "required": [],
      "additionalProperties": false
    },
    "strict": true
  }
}

The resulting response using the OpenAI Python client is:

ChatCompletion(
  choices=[
    Choice(
      finish_reason='stop',
      message=ChatCompletionMessage(
        content='',
        role='assistant',
        function_call=None,
        tool_calls=None
      )
    )
  ]
)

finish_reason: stop
content: ''
tool_calls: None
reasoning: None

Usage for one of the failing requests:

completion_tokens=19
prompt_tokens=63
total_tokens=82

Therefore the significant A/B difference appears to be:

Tool without "strict": true
        |
        +--> finish_reason = tool_calls
             tool call generated successfully

Tool with "strict": true
        |
        +--> finish_reason = stop
             content = ""
             tool_calls = None

Expected behavior

Adding:

"strict": true

to an otherwise working OpenAI-compatible function definition should
not prevent the model from making the tool call.

I would expect a response equivalent to:

{
  "finish_reason": "tool_calls",
  "message": {
    "tool_calls": [
      {
        "type": "function",
        "function": {
          "name": "transfer_to_customer_retention_agent",
          "arguments": "{}"
        }
      }
    ]
  }
}

The function arguments in this example already conform to the supplied
JSON schema.


Logs

LocalAI starts the model with:

INFO BackendLoader starting
modelID="qwen3.5-9b-glm5.1-distill-v1"
backend="llama-cpp"
model="llama-cpp/models/qwen3.5-9b-glm5.1-distill-v1/Qwen3.5-9B-GLM5.1-Distill-v1-Q4_K_M.gguf"

The effective runtime configuration is:

context=2048
n_batch=512
n_gpu_layers=0
parallel="1"
flash_attention="auto"
f16=false

During failing requests LocalAI repeatedly reports:

WARN Backend returned empty response, retrying attempt=1 maxRetries=5
WARN Backend returned empty response, retrying attempt=2 maxRetries=5
WARN Backend returned empty response, retrying attempt=3 maxRetries=5
WARN Backend returned empty response, retrying attempt=4 maxRetries=5
WARN Backend returned empty response, retrying attempt=5 maxRetries=5
WARN Backend produced reasoning without actionable content, retrying reasoning_len=0 attempt=6

The HTTP request nevertheless completes with:

POST /v1/chat/completions status=200

LocalAI version from the startup log:

INFO Using forced capability from environment variable
capability="nvidia"
env="LOCALAI_FORCE_META_BACKEND_CAPABILITY"

INFO LocalAI version
version="v4.9.0 (f7ad3f70eb5d8a0ddf80e08557f0d7df28cf032e)"

Additional context

The problem was originally discovered while testing LocalAI with the
OpenAI Agents Python SDK version 0.22.0.

An Agents SDK handoff is exposed to the model as a function tool similar
to:

Tool name:
transfer_to_customer_retention_agent

Schema:
{
  "additionalProperties": false,
  "type": "object",
  "properties": {},
  "required": []
}

Strict:
True

When this is sent through LocalAI, the Agents SDK receives an empty
model output and therefore cannot perform the handoff:

ModelResponse(
    output=[],
    usage=Usage(
        requests=1,
        input_tokens=100,
        output_tokens=23,
        total_tokens=123
    )
)

As a workaround I replaced the native handoff with an ordinary Agents
SDK function tool declared with:

@function_tool(strict_mode=False)
def request_handoff(target_agent: str, reason: str) -> str:
    ...

The same Qwen model then successfully determines that a handoff is
required and calls the tool:

[TOOL] request_handoff called
[TOOL] target_agent = Customer Retention Agent
[TOOL] reason = Customer wants to cancel their order and account after experiencing repeated order delays.

Python can then perform the actual agent transfer.

This suggests that the model is capable of the required tool selection
and function call, and that the failure is specifically associated with
the strict: true tool definition.

Summary of observed behaviour:

Test Result
Normal chat completion Works
Function tool without strict Works
Function tool with "strict": true Fails with empty response
OpenAI Agents SDK native handoff Fails
Agents SDK strict_mode=False function tool Works
主要言語
Go
スター
49.2k
フォーク
4.5k
平均マージ
1日 7時間
マージ済み PR(30日)
340

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

mudler/LocalAI のほかの issue

mudler/LocalAI の issue をすべて見る

似ている issue

Go の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。