Optional EvalPort adapter for Trajectory/TrajectoryGroup (portable eval-result interchange)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 42/100
Research direction
Start with src/art/trajectories/init.py and the EvalPort SPEC.md, then review examples/tic_tac_toe/rollout.py and examples/mcp-rl/ for existing rollout flows. Done means an agreed standalone optional adapter location and a tested bidirectional mapping between Trajectory/TrajectoryGroup and EvalPort results without changing the training loop or requiring a dependency.
Written by the indexing model from the issue text.
Description
I've been reading src/art/trajectories/__init__.py — Trajectory (messages_and_choices, reward: float, metrics: dict[str, float | int | bool]) and TrajectoryGroup (grouping multiple rollouts of the same task, with its own metrics). That's a clean shape for "one rollout's outcome plus its scalar reward and side metrics" — closer to an eval result than most RL trajectory formats I've seen, since metrics already separates named auxiliary signals from the training-facing reward.
I maintain EvalPort, a JSON-Schema-based interchange spec (TestCase/Suite/Grader/Result/ResultSet) for portable LLM eval data, so eval/grading data isn't locked to one framework's format. It's early-stage (~35 shipped adapters, no notable star count — being upfront about that).
The mapping here is fairly direct: Trajectory.reward → EvalPort GraderResult.score, Trajectory.metrics → additional named GraderResults (the same "multiple named signals per outcome" pattern EvalPort's schema is built around), messages_and_choices → the Result transcript, and a TrajectoryGroup (multiple rollouts of one task, as used for GRPO's relative comparisons) → an EvalPort ResultSet grouped by task. That would let an ART training/eval run's rollouts be read by grading tooling built for other frameworks, or let a suite of tasks authored as portable EvalPort TestCases drive ART rollouts instead of a one-off rollout.py per project (as in examples/tic_tac_toe/rollout.py, examples/mcp-rl/, etc.).
I'd propose a standalone, optional adapter doing that conversion in both directions. No required dependency, no change to Trajectory/TrajectoryGroup or the training loop.
Happy to build this as a PR into ART (e.g. src/art/interop/evalport.py), or as a standalone package in EvalPort's own adapters/ directory with zero footprint on this repo — whichever you'd prefer. Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md
— Sahi, independent contributor (not affiliated with OpenPipe)
- Dominant language
- Python
- Stars
- 10.8k
- Forks
- 997
- Avg merge
- 10h 1m
- Merged PRs (30d)
- 117
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from OpenPipe/ART
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 54/100
OpenPipe/ART#961 · 3 comments ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
OpenPipe/ART#949 · 5 comments ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 10/100
Maintainers usually reply within 1 day
-
enhancement
Difficulty 4/5 3-5 days Newbie friendliness 48/100
Maintainers usually reply within 1 day
Similar issues
-
adr
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
kristofdegrave/homeassistant-smart-charging#1607 ·
Maintainers usually reply within 1 day
-
namespace operations
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
EclipseFdn/open-vsx.org#13665 ·
Maintainers usually reply within 1 day
-
doc good first issue help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
collective/icalendar#1865 · 2 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
canonical/opentelemetry-collector-operator#409 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 85/100
mozilla/addons-release-tests#1243 ·
Maintainers usually reply within 1 day