RULER + Litellm with Ollama: `ollama_chat` recommended, but `/api/chat` cannot generate required JSON schema
Maintainer thường phản hồi trong vòng 3 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 35/100
Hướng nghiên cứu
Không có tệp, bài kiểm thử hoặc điểm vào nào được nêu. Hãy xác định cấu hình judge LiteLLM/Ollama của RULER và tái hiện lỗi xác thực schema với ollama_chat; hoàn thành có nghĩa là либо cấu hình được hỗ trợ cùng các giới hạn của nó được ghi lại, либо hỗ trợ structured output hoạt động cho /api/chat.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Description
When configuring Ollama via Litellm, the Litellm docs recommend using:
ollama_chatfor better responses
This uses Ollama’s /api/chat endpoint.
However, in ART’s RULER integration, the judge model needs to:
- Produce JSON-structured outputs (e.g. match a Pydantic schema / tool schema) so RULER can parse correctness, reasoning, etc.
The problem:
-
With Ollama’s
/api/chat(viaollama_chat), the responses are not reliably:- Valid JSON matching the expected schema, or
- Tool-call-compatible in the way RULER expects.
-
As a result:
- The judge’s responses often fail schema validation.
- RULER cannot extract the required fields, and the scoring fails.
Because of that, I’m forced to fall back to the older /api/generate endpoint (ollama), even though:
- Litellm explicitly recommends
ollama_chatoverollama. /api/chatis the more modern endpoint.
What I expect
-
Either:
- Official support / documentation for using Ollama’s
/api/chat(ollama_chat) as a RULER judge with JSON-schema / tool-call style responses; or - Clear guidance that for RULER’s JSON-schema needs, we must currently use
/api/generateandollama.
- Official support / documentation for using Ollama’s
What actually happens
-
Using
ollama_chat:- The
/api/chatendpoint does not produce the JSON / tool-call schema RULER expects. - Judge calls fail schema validation.
- The
-
Using
ollama:- RULER works better, but we lose out on the newer
/api/chatbehavior Litellm recommends.
- RULER works better, but we lose out on the newer
Request
-
Please:
-
Document the supported Ollama configuration for RULER (which engine, which endpoint, any special settings).
-
If possible, add direct support for
ollama_chat+/api/chatthat:- Ensures JSON-mode / tool-call style output works with RULER’s structured response expectations.
-
Or provide example configs + templates that make Ollama’s
/api/chatusable with RULER.
-
- Ngôn ngữ chính
- Python
- Star
- 10.8k
- Fork
- 989
- Merge trung bình
- 11 giờ 38 phút
- Pull request đã merge (30 ngày)
- 104
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của OpenPipe/ART
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 54/100
OpenPipe/ART#961 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
OpenPipe/ART#949 · 5 bình luận ·
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 10/100
Maintainer thường phản hồi trong vòng 3 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 42/100
Maintainer thường phản hồi trong vòng 3 ngày
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
solana-foundation/pay-kit#341 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
nasa/python_cmr#123 ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
EleutherAI/lm-evaluation-harness#4243 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area: dashboard bug perceived difficulty: 3
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Nitjsefnie-Harness-Commons/daedalus#1179 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
cusp-ai-oss/tojax#17 ·