Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Optional EvalPort adapter for Trajectory/TrajectoryGroup (portable eval-result interchange)

Open
#935 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
42/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
python
Domain
data, tooling

Research direction

Start with src/art/trajectories/init.py and the EvalPort SPEC.md, then review examples/tic_tac_toe/rollout.py and examples/mcp-rl/ for existing rollout flows. Done means an agreed standalone optional adapter location and a tested bidirectional mapping between Trajectory/TrajectoryGroup and EvalPort results without changing the training loop or requiring a dependency.

Written by the indexing model from the issue text.

Description

I've been reading src/art/trajectories/__init__.py — Trajectory (messages_and_choices, reward: float, metrics: dict[str, float | int | bool]) and TrajectoryGroup (grouping multiple rollouts of the same task, with its own metrics). That's a clean shape for "one rollout's outcome plus its scalar reward and side metrics" — closer to an eval result than most RL trajectory formats I've seen, since metrics already separates named auxiliary signals from the training-facing reward.

I maintain EvalPort, a JSON-Schema-based interchange spec (TestCase/Suite/Grader/Result/ResultSet) for portable LLM eval data, so eval/grading data isn't locked to one framework's format. It's early-stage (~35 shipped adapters, no notable star count — being upfront about that).

The mapping here is fairly direct: Trajectory.reward → EvalPort GraderResult.score, Trajectory.metrics → additional named GraderResults (the same "multiple named signals per outcome" pattern EvalPort's schema is built around), messages_and_choices → the Result transcript, and a TrajectoryGroup (multiple rollouts of one task, as used for GRPO's relative comparisons) → an EvalPort ResultSet grouped by task. That would let an ART training/eval run's rollouts be read by grading tooling built for other frameworks, or let a suite of tasks authored as portable EvalPort TestCases drive ART rollouts instead of a one-off rollout.py per project (as in examples/tic_tac_toe/rollout.py, examples/mcp-rl/, etc.).

I'd propose a standalone, optional adapter doing that conversion in both directions. No required dependency, no change to Trajectory/TrajectoryGroup or the training loop.

Happy to build this as a PR into ART (e.g. src/art/interop/evalport.py), or as a standalone package in EvalPort's own adapters/ directory with zero footprint on this repo — whichever you'd prefer. Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md

— Sahi, independent contributor (not affiliated with OpenPipe)

Dominant language
Python
Stars
10.8k
Forks
997
Avg merge
10h 1m
Merged PRs (30d)
117

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from OpenPipe/ART

All issues in OpenPipe/ART

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.