RULER scoring + training tightly coupled to Litellm/OpenAI, cannot cleanly use ChatOllama/ChatNVIDIA as judge/inference models
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- ollama, python
- Domain
- ai, machine-learning
Research direction
Start with ruler_score_group and its related helpers, then inspect init_chat_model and the separate issue it references. Trace where LiteLLM and OpenAI-style assumptions enter RULER and training; done means the project either supports LangChain BaseChatModel providers such as ChatOllama and ChatNVIDIA or documents an arbitrary judge_fn integration path.
Written by the indexing model from the issue text.
Description
Description
In my setup, I want to:
-
Use a local Ollama server (and potentially NVIDIA’s API in the future) as:
- The main agent model (for rollouts).
- The judge model (for RULER scoring).
However, the current ART stack makes this very difficult because:
- RULER scoring (
ruler_score_groupand related helpers) rely on Litellm in a way that expects OpenAI-style models. init_chat_modelalso wraps everything in aChatOpenAIinstance (see separate issue).- This means I cannot simply pass
ChatOllamaorChatNVIDIA(LangChain chat models) as the inference/judge model for training.
Practically:
-
If I try to step away from OpenAI and use:
- Local Ollama for inference
- Non-OpenAI providers as judges
-
I run into incompatibilities where:
- RULER expects Litellm’s OpenAI-style model identifiers and behavior.
- ART’s helpers are “too bound” to OpenAI semantics.
What I’d like
-
A more provider-agnostic design for:
- RULER scoring
- Training
init_chat_model
-
The ability to cleanly use:
ChatOllama(LangChain)ChatNVIDIA- or other LangChain
BaseChatModelimplementations
-
Without having to hack around Litellm / OpenAI assumptions.
Why this matters
-
ART is otherwise a great framework for agent RL.
-
Many users want to move to:
- Local models (Ollama)
- Different clouds (NVIDIA, etc.)
-
Tight coupling to OpenAI via Litellm in the RULER path makes this significantly harder.
Request
-
Please consider:
- Abstracting RULER to accept any LangChain-compatible
ChatModelfor structured scoring. - Or providing a documented way to plug in non-OpenAI judgment models (e.g. a “judge_fn” that uses arbitrary models).
- Abstracting RULER to accept any LangChain-compatible
- Dominant language
- Python
- Stars
- 10.8k
- Forks
- 997
- Avg merge
- 10h 1m
- Merged PRs (30d)
- 117
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from OpenPipe/ART
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 54/100
OpenPipe/ART#961 · 3 comments ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
OpenPipe/ART#949 · 5 comments ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 10/100
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 42/100
Maintainers usually reply within 1 day
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
mozilla/bedrock#17413 · 1 reaction ·
Maintainers usually reply within 2 days
-
instance instance add
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
searxng/searx-instances#943 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
bug tools
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
lance-format/lance#9655 ·
Maintainers usually reply within 2 days