Docker Model Runner (llama.cpp backend) returns HTTP 500 "does not match the expected peg-native format" when a single chat turn requires 2+ sequential tool/function calls
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 68/100
Hướng nghiên cứu
Bắt đầu bằng cách tái hiện yêu cầu đối với endpoint /engines/llama.cpp/v1/chat/completions với hai lần gọi công cụ, sau đó kiểm tra log của model Docker và tệp đính kèm DMR_BUG_REPORT_logs.txt để tìm cảnh báo phân tích cú pháp peg-native. Truy vết đường dẫn phân tích cú pháp nhiều lần gọi và làm cho các lần gọi công cụ tuần tự thành công, hoặc trả về một lỗi riêng biệt đã được tài liệu hóa khi không được hỗ trợ; xác minh cả việc tái hiện lẫn hành vi hiện có của lần gọi đơn.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Bug report: Docker Model Runner (llama.cpp backend) returns HTTP 500 "does not match the expected peg-native format" when a single chat turn requires 2+ sequential tool/function calls
File this at: https://github.com/docker/model-runner/issues
Attach alongside this report: DMR_BUG_REPORT_logs_raw.txt (raw docker model logs excerpts, same folder)
Summary
When a chat completion request against the llama.cpp backend (OpenAI-compatible
/engines/llama.cpp/v1/chat/completions endpoint) results in the model producing
more than one tool/function call in a single assistant turn, the backend's
structured-output ("peg-native") parser fails to parse anything past the first
call and the request fails with HTTP 500:
{"error":{"code":500,"message":"The model produced output that does not match the expected peg-native format","type":"server_error"}}
The corresponding server log always shows the unparsed raw output containing
multiple {"name": ..., "parameters": {...}} objects concatenated with ;
— i.e. the model DID produce a reasonable, syntactically-valid multi-call
response, but the peg-native grammar/parser only expects (and can only handle)
a single call per turn.
This is fully reproducible, was observed identically across 3 independent
testing sessions on our end, and blocks any workflow where an agent needs to
call 2+ tools in the same turn (e.g. "create these two files" or "do A, B, and
C" in one instruction) — a common, realistic agent usage pattern.
Environment
- Docker Desktop / Docker Engine: 29.8.0
- Docker Model Runner CLI: Client v1.2.6, Server v1.2.8 (Engine: Docker Desktop)
- Backend:
llama.cpp b9879-cuda(sha256:57130b8af4f80d754ee93e8855e74a9bc6c13e07f99b9becab720cce2658199e) - Model:
ai/llama3.1:8B-Q4_K_M(also reproduces on plainai/llama3.1:8breference) - GPU: NVIDIA GeForce RTX 4060 (8GB VRAM), driver 32.0.16.1692
- OS: Windows 11, Docker Desktop with WSL2 backend
- Endpoint under test:
http://localhost:12434/engines/llama.cpp/v1/chat/completions
(OpenAI-compatible chat completions,toolsparam populated,stream: false)
Steps to reproduce
-
Start DMR with the llama.cpp backend and load
ai/llama3.1:8B-Q4_K_M(or
any llama3.1 8B variant). -
Send a chat completion request with 2+ tool schemas available (e.g.
write_file,read_file,list_directory— simple JSON-schema function
defs, nothing exotic) and a user instruction that clearly asks for two
distinct tool calls in one turn, e.g.:"Create a file at reports/summary.md containing exactly: 'Daily summary:
all tasks complete.' Then create a second file at logs/log1.txt
containing exactly: 'Log entry 1: system start.' Use your write_file
tool for both, one at a time." -
Observe: the request returns HTTP 500 with the message above.
-
Check
docker model logsfor the same request — it will show the model's
raw (correct, well-formed) multi-call output, immediately followed by the
peg-native parse failure.
Reproduced 7+ times across 3 separate sessions on our end, always in
response to a turn requiring 2+ tool calls; never observed on single-tool-call
turns.
Actual raw log evidence (representative example)
W common_chat_peg_parse: unparsed peg-native output: ; {"name": "write_file", "parameters": {"path": "notes/b.txt", "content": "Beta"}}; {"name": "list_directory", "parameters": {"path": "notes"}}
W srv operator (): got exception: {"error":{"code":500,"message":"The model produced output that does not match the expected peg-native format","type":"server_error"}}
Note the raw output the model produced is two syntactically valid JSON tool-call
objects separated by ; — this is a parser limitation (single-call-only
grammar), not a malformed-model-output problem.
Additional representative examples (all same signature, different tool
args) are included in the attached DMR_BUG_REPORT_logs_raw.txt.
Expected behavior
The peg-native/structured-output parser should support parsing multiple
sequential tool calls emitted in a single assistant turn (as OpenAI's own API
and Ollama's /api/chat both do — see comparison below), OR, if multi-call
support is intentionally out of scope for this backend, the server should
document this constraint and/or return a clearer, distinct error/status
indicating "multiple tool calls not supported" rather than a generic 500
implying malformed output.
Comparison: identical prompt/tools against Ollama (llama3.1:8b, /api/chat)
Run side-by-side with identical tool schemas and prompts, Ollama's /api/chat
correctly parses and executes both tool calls from the same "create two files"
style instruction, with zero errors, across every run tested (multiple runs
across multiple sessions). This strongly suggests the issue is specific to
DMR's llama.cpp peg-native parser's single-call assumption, not a fundamental
limitation of the underlying llama3.1:8b model's tool-calling ability.
Impact
Any agent workflow that asks for 2+ actions in a single turn (very common:
"do X and then Y", "make these N files", batch instructions) will hard-fail
with a 500 on this backend, even though the underlying model handles the
request correctly. This was a decision-relevant finding in our own
evaluation of DMR vs. Ollama as a default local model runtime, since our
planned usage pattern (an AI assistant/agent doing file/calendar/contact
operations) routinely issues multi-tool-call turns.
Secondary, related-but-distinct observation (optional, included for completeness)
During the same testing, we also observed a separate HTTP 500 caused by a
JSON parse error when a prior assistant turn contained a non-ASCII character
(specifically "é" in the word "décor"), which came back mangled
(ill-formed UTF-8 byte) when re-serialized into a later request in the same
conversation:
{"error":{"code":500,"message":"[json.exception.parse_error.101] parse error at line 28, column 224: syntax error while parsing value - invalid string: ill-formed UTF-8 byte; last read: '\"A most refined and elegant choice, Vince, sir. Deep green is a color that evokes a sense of sophistication and refinement. I shall take note of your preference and see to it that the déc'","type":"server_error"}}
We believe this is a distinct bug from the peg-native multi-call issue
above (different trigger: non-ASCII character in conversation history, not
multiple tool calls) and are including it here only as a secondary data point
in case it's useful to a maintainer — happy to file this as a separate issue
if preferred.
- Ngôn ngữ chính
- Go
- Star
- 656
- Fork
- 156
- Merge trung bình
- 3 ngày 10 giờ
- Pull request đã merge (30 ngày)
- 8
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Không có mẫu pull request
- Không có hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của docker/model-runner
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
docker/model-runner#1077 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
docker/model-runner#1073 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 42/100
docker/model-runner#1059 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
docker/model-runner#1058 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
docker/model-runner#1056 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của docker/model-runner
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
war-and-code/dircue#200 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
actor/human kind/bug priority/important-soon triage-accepted
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 66/100
kelos-dev/kelos#1804 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày
-
`date` → `date-time` (and `time` → `date-time`) is classified as a widening, but a date is not a valid date-timeCó thể đã có người làm @reuvenharrison đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày
-
fix: invalid GPU spec in SparkApplication is silently ignoredCó thể đã có người làm @pratik-naik003 đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
kubeflow/spark-operator#3223 ·
Maintainer thường phản hồi trong vòng 5 ngày