Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Docker Model Runner (llama.cpp backend) returns HTTP 500 "does not match the expected peg-native format" when a single chat turn requires 2+ sequential tool/function calls

Đang mở
#1,063 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
68/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
docker, go
Lĩnh vực
ai, api, backend

Hướng nghiên cứu

Bắt đầu bằng cách tái hiện yêu cầu đối với endpoint /engines/llama.cpp/v1/chat/completions với hai lần gọi công cụ, sau đó kiểm tra log của model Docker và tệp đính kèm DMR_BUG_REPORT_logs.txt để tìm cảnh báo phân tích cú pháp peg-native. Truy vết đường dẫn phân tích cú pháp nhiều lần gọi và làm cho các lần gọi công cụ tuần tự thành công, hoặc trả về một lỗi riêng biệt đã được tài liệu hóa khi không được hỗ trợ; xác minh cả việc tái hiện lẫn hành vi hiện có của lần gọi đơn.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Bug report: Docker Model Runner (llama.cpp backend) returns HTTP 500 "does not match the expected peg-native format" when a single chat turn requires 2+ sequential tool/function calls

File this at: https://github.com/docker/model-runner/issues
Attach alongside this report: DMR_BUG_REPORT_logs_raw.txt (raw docker model logs excerpts, same folder)


Summary

When a chat completion request against the llama.cpp backend (OpenAI-compatible
/engines/llama.cpp/v1/chat/completions endpoint) results in the model producing
more than one tool/function call in a single assistant turn, the backend's
structured-output ("peg-native") parser fails to parse anything past the first
call and the request fails with HTTP 500:

{"error":{"code":500,"message":"The model produced output that does not match the expected peg-native format","type":"server_error"}}

The corresponding server log always shows the unparsed raw output containing
multiple {"name": ..., "parameters": {...}} objects concatenated with ;
— i.e. the model DID produce a reasonable, syntactically-valid multi-call
response, but the peg-native grammar/parser only expects (and can only handle)
a single call per turn.

This is fully reproducible, was observed identically across 3 independent
testing sessions on our end, and blocks any workflow where an agent needs to
call 2+ tools in the same turn (e.g. "create these two files" or "do A, B, and
C" in one instruction) — a common, realistic agent usage pattern.


Environment

  • Docker Desktop / Docker Engine: 29.8.0
  • Docker Model Runner CLI: Client v1.2.6, Server v1.2.8 (Engine: Docker Desktop)
  • Backend: llama.cpp b9879-cuda (sha256:57130b8af4f80d754ee93e8855e74a9bc6c13e07f99b9becab720cce2658199e)
  • Model: ai/llama3.1:8B-Q4_K_M (also reproduces on plain ai/llama3.1:8b reference)
  • GPU: NVIDIA GeForce RTX 4060 (8GB VRAM), driver 32.0.16.1692
  • OS: Windows 11, Docker Desktop with WSL2 backend
  • Endpoint under test: http://localhost:12434/engines/llama.cpp/v1/chat/completions
    (OpenAI-compatible chat completions, tools param populated, stream: false)

Steps to reproduce

  1. Start DMR with the llama.cpp backend and load ai/llama3.1:8B-Q4_K_M (or
    any llama3.1 8B variant).

  2. Send a chat completion request with 2+ tool schemas available (e.g.
    write_file, read_file, list_directory — simple JSON-schema function
    defs, nothing exotic) and a user instruction that clearly asks for two
    distinct tool calls in one turn
    , e.g.:

    "Create a file at reports/summary.md containing exactly: 'Daily summary:
    all tasks complete.' Then create a second file at logs/log1.txt
    containing exactly: 'Log entry 1: system start.' Use your write_file
    tool for both, one at a time."

  3. Observe: the request returns HTTP 500 with the message above.

  4. Check docker model logs for the same request — it will show the model's
    raw (correct, well-formed) multi-call output, immediately followed by the
    peg-native parse failure.

Reproduced 7+ times across 3 separate sessions on our end, always in
response to a turn requiring 2+ tool calls; never observed on single-tool-call
turns.


Actual raw log evidence (representative example)

W common_chat_peg_parse: unparsed peg-native output: ; {"name": "write_file", "parameters": {"path": "notes/b.txt", "content": "Beta"}}; {"name": "list_directory", "parameters": {"path": "notes"}}
W srv   operator (): got exception: {"error":{"code":500,"message":"The model produced output that does not match the expected peg-native format","type":"server_error"}}

Note the raw output the model produced is two syntactically valid JSON tool-call
objects separated by ; — this is a parser limitation (single-call-only
grammar), not a malformed-model-output problem.

Additional representative examples (all same signature, different tool
args) are included in the attached DMR_BUG_REPORT_logs_raw.txt.


Expected behavior

The peg-native/structured-output parser should support parsing multiple
sequential tool calls emitted in a single assistant turn (as OpenAI's own API
and Ollama's /api/chat both do — see comparison below), OR, if multi-call
support is intentionally out of scope for this backend, the server should
document this constraint and/or return a clearer, distinct error/status
indicating "multiple tool calls not supported" rather than a generic 500
implying malformed output.


Comparison: identical prompt/tools against Ollama (llama3.1:8b, /api/chat)

Run side-by-side with identical tool schemas and prompts, Ollama's /api/chat
correctly parses and executes both tool calls from the same "create two files"
style instruction, with zero errors, across every run tested (multiple runs
across multiple sessions). This strongly suggests the issue is specific to
DMR's llama.cpp peg-native parser's single-call assumption, not a fundamental
limitation of the underlying llama3.1:8b model's tool-calling ability.


Impact

Any agent workflow that asks for 2+ actions in a single turn (very common:
"do X and then Y", "make these N files", batch instructions) will hard-fail
with a 500 on this backend, even though the underlying model handles the
request correctly. This was a decision-relevant finding in our own
evaluation of DMR vs. Ollama as a default local model runtime, since our
planned usage pattern (an AI assistant/agent doing file/calendar/contact
operations) routinely issues multi-tool-call turns.


Secondary, related-but-distinct observation (optional, included for completeness)

During the same testing, we also observed a separate HTTP 500 caused by a
JSON parse error when a prior assistant turn contained a non-ASCII character
(specifically "é" in the word "décor"), which came back mangled
(ill-formed UTF-8 byte) when re-serialized into a later request in the same
conversation:

{"error":{"code":500,"message":"[json.exception.parse_error.101] parse error at line 28, column 224: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"A most refined and elegant choice, Vince, sir. Deep green is a color that evokes a sense of sophistication and refinement. I shall take note of your preference and see to it that the déc'","type":"server_error"}}

We believe this is a distinct bug from the peg-native multi-call issue
above (different trigger: non-ASCII character in conversation history, not
multiple tool calls) and are including it here only as a secondary data point
in case it's useful to a maintainer — happy to file this as a separate issue
if preferred.

DMR_BUG_REPORT_logs.txt

Ngôn ngữ chính
Go
Star
656
Fork
156
Merge trung bình
3 ngày 10 giờ
Pull request đã merge (30 ngày)
8

Chuẩn bị môi trường

  • Có Dockerfile hoặc tệp Docker Compose
  • Không có mẫu pull request
  • Không có hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của docker/model-runner

Tất cả issue của docker/model-runner

Issue tương tự

Thêm issue về Go

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.