Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Optional EvalPort adapter for Trajectory/TrajectoryGroup (portable eval-result interchange)

オープン
#935 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 3 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
42/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
活発
技術スタック
python
領域
data, tooling

調査の方向性

Start with src/art/trajectories/init.py and the EvalPort SPEC.md, then review examples/tic_tac_toe/rollout.py and examples/mcp-rl/ for existing rollout flows. Done means an agreed standalone optional adapter location and a tested bidirectional mapping between Trajectory/TrajectoryGroup and EvalPort results without changing the training loop or requiring a dependency.

索引モデルが issue の本文から書いたものです。

説明

I've been reading src/art/trajectories/__init__.py — Trajectory (messages_and_choices, reward: float, metrics: dict[str, float | int | bool]) and TrajectoryGroup (grouping multiple rollouts of the same task, with its own metrics). That's a clean shape for "one rollout's outcome plus its scalar reward and side metrics" — closer to an eval result than most RL trajectory formats I've seen, since metrics already separates named auxiliary signals from the training-facing reward.

I maintain EvalPort, a JSON-Schema-based interchange spec (TestCase/Suite/Grader/Result/ResultSet) for portable LLM eval data, so eval/grading data isn't locked to one framework's format. It's early-stage (~35 shipped adapters, no notable star count — being upfront about that).

The mapping here is fairly direct: Trajectory.reward → EvalPort GraderResult.score, Trajectory.metrics → additional named GraderResults (the same "multiple named signals per outcome" pattern EvalPort's schema is built around), messages_and_choices → the Result transcript, and a TrajectoryGroup (multiple rollouts of one task, as used for GRPO's relative comparisons) → an EvalPort ResultSet grouped by task. That would let an ART training/eval run's rollouts be read by grading tooling built for other frameworks, or let a suite of tasks authored as portable EvalPort TestCases drive ART rollouts instead of a one-off rollout.py per project (as in examples/tic_tac_toe/rollout.py, examples/mcp-rl/, etc.).

I'd propose a standalone, optional adapter doing that conversion in both directions. No required dependency, no change to Trajectory/TrajectoryGroup or the training loop.

Happy to build this as a PR into ART (e.g. src/art/interop/evalport.py), or as a standalone package in EvalPort's own adapters/ directory with zero footprint on this repo — whichever you'd prefer. Spec: https://github.com/adhabnr-ux/evalport/blob/main/SPEC.md

— Sahi, independent contributor (not affiliated with OpenPipe)

主要言語
Python
スター
10.8k
フォーク
989
平均マージ
11時間 38分
マージ済み PR(30日)
104

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

OpenPipe/ART のほかの issue

OpenPipe/ART の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。