ADF judge: bounded decision-model agreement experiment (current judge vs Jev vs Levanto Sage vs local)
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 30/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- rust
- 領域
- ai, testing-qa
調査の方向性
Start with the claude-api build-eval discipline and trace the current ADF judge plus the llm-proxy typesafe::/jev:: adapter. Check how a backend using the /v1/systemone shape could be invoked, then define the sampled cases, permutations, duplicate runs, noise floor, Wilson CIs, and cost-per-task measurements. Done means reporting agreement, cost, permutation swing, and a keep/trial/drop recommendation for every arm.
索引モデルが issue の本文から書いたものです。
説明
Goal
A bounded experiment on one ADF judge decision: does a typed decision model agree with the current reasoned verdict, and what does each option cost per completed task (not per call)?
Arms
- The current ADF judge verdict (baseline).
- TypeSafe Jev
Choice(already reachable through the llm-proxytypesafe::/jev::adapter). - Levanto Sage
choice/yesno, on its free tier (100 decisions a month), withreasoning: "off"and"on"as separate conditions. - A local open-weights backend (openjev-sglang or AnyJev), if one is available on the same
/v1/systemoneshape.
Method
Build the eval with the claude-api build-eval discipline:
- Cases sampled from real judge inputs; hard cases chosen by a human, not by current failures.
- Permute the option order for every decision-model arm. A permutation swing larger than the arm-to-arm difference invalidates that arm's result.
- Run each grader twice to measure its flip rate.
- State the noise floor before comparing arms. Report agreement with Wilson CIs, not point estimates.
Out of scope
Replacing the reasoned judge. Neither vendor returns a justification: Sage's meta.reasoning is telemetry only. So a decision-model verdict cannot satisfy the requirement that verdicts be auditable. This experiment measures agreement and cost only.
Acceptance
- Agreement and cost per task for each arm, with CIs and the permutation swing.
- A written recommendation: keep, trial, or drop.
Related
- Lane-policy decision for metered classifiers (separate issue).
Mirror of Gitea terraphim/terraphim-ai#3418 (source of truth).
- 主要言語
- Rust
- スター
- 65
- フォーク
- 5
- 平均マージ
- 1時間 17分
- マージ済み PR(30日)
- 2
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
terraphim/terraphim-ai のほかの issue
-
terraphim_rlm: require a per-sub-result verification plan (recheck / control / independent route)オープン
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
terraphim/terraphim-ai#969 · コメント 1 件 ·
-
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
terraphim/terraphim-ai#968 · コメント 1 件 ·
-
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
terraphim/terraphim-ai#967 ·
-
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
terraphim/terraphim-ai#966 ·
-
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
terraphim/terraphim-ai#885 ·
terraphim/terraphim-ai の issue をすべて見る
似ている issue
-
agent/sec-check hive/hive-school-tunaos security
難易度 2/5 1〜3時間 初心者へのやさしさ 90/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 90/100
Sovereign-Labs/sovereign-sdk#3070 ·
-
Accepts Invalid URLオープンbug
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
メンテナーはふだん 1 日以内に返信
-
area:prove bug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
Chelis-Lang/chelis#3317 ·
メンテナーはふだん 1 日以内に返信
-
[macOS Desktop] New sidebar hover navigation accidentally switches sections while reaching a chatオープンapp bug
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
メンテナーはふだん 1 日以内に返信