ADF judge: bounded decision-model agreement experiment (current judge vs Jev vs Levanto Sage vs local)
還沒有人認領這個 Issue。
評估
- 難度
- 5/5
- 預估耗時
- 一週以上
- 新手友好度
- 30/100
- Issue 類型
- 功能
- 描述清晰度
- 基本清楚
- 活躍度
- 活躍
- 技術堆疊
- rust
- 領域
- ai, testing-qa
研究方向
Start with the claude-api build-eval discipline and trace the current ADF judge plus the llm-proxy typesafe::/jev:: adapter. Check how a backend using the /v1/systemone shape could be invoked, then define the sampled cases, permutations, duplicate runs, noise floor, Wilson CIs, and cost-per-task measurements. Done means reporting agreement, cost, permutation swing, and a keep/trial/drop recommendation for every arm.
由索引模型根據 Issue 內容生成。
描述
Goal
A bounded experiment on one ADF judge decision: does a typed decision model agree with the current reasoned verdict, and what does each option cost per completed task (not per call)?
Arms
- The current ADF judge verdict (baseline).
- TypeSafe Jev
Choice(already reachable through the llm-proxytypesafe::/jev::adapter). - Levanto Sage
choice/yesno, on its free tier (100 decisions a month), withreasoning: "off"and"on"as separate conditions. - A local open-weights backend (openjev-sglang or AnyJev), if one is available on the same
/v1/systemoneshape.
Method
Build the eval with the claude-api build-eval discipline:
- Cases sampled from real judge inputs; hard cases chosen by a human, not by current failures.
- Permute the option order for every decision-model arm. A permutation swing larger than the arm-to-arm difference invalidates that arm's result.
- Run each grader twice to measure its flip rate.
- State the noise floor before comparing arms. Report agreement with Wilson CIs, not point estimates.
Out of scope
Replacing the reasoned judge. Neither vendor returns a justification: Sage's meta.reasoning is telemetry only. So a decision-model verdict cannot satisfy the requirement that verdicts be auditable. This experiment measures agreement and cost only.
Acceptance
- Agreement and cost per task for each arm, with CIs and the permutation swing.
- A written recommendation: keep, trial, or drop.
Related
- Lane-policy decision for metered classifiers (separate issue).
Mirror of Gitea terraphim/terraphim-ai#3418 (source of truth).
- 主要語言
- Rust
- 星號
- 65
- 分支
- 5
- 平均合併
- 1 小時 17 分鐘
- 30 天內合併 PR
- 2
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
terraphim/terraphim-ai 的其他 Issue
-
terraphim_rlm: require a per-sub-result verification plan (recheck / control / independent route)未關閉
難度 4/5 3-5 天 新手友好度 48/100
terraphim/terraphim-ai#969 · 1 則留言 ·
-
難度 5/5 一週以上 新手友好度 35/100
terraphim/terraphim-ai#968 · 1 則留言 ·
-
難度 5/5 一週以上 新手友好度 35/100
terraphim/terraphim-ai#967 ·
-
難度 4/5 3-5 天 新手友好度 48/100
terraphim/terraphim-ai#966 ·
-
難度 5/5 一週以上 新手友好度 25/100
terraphim/terraphim-ai#885 ·
查看 terraphim/terraphim-ai 的全部 Issue
相似的 Issue
-
✨ enhancement needs-discussion
難度 1/5 1 小時以內 新手友好度 85/100
-
area:docs documentation good first issue priority:low
難度 2/5 1-3 小時 新手友好度 68/100
維護者通常 1 天內回覆
-
triage:accepted
難度 2/5 1-3 小時 新手友好度 65/100
open-telemetry/otel-arrow#4343 ·
維護者通常 2 天內回覆
-
bug
難度 2/5 1-3 小時 新手友好度 75/100
mishraprafful/multihull#150 ·
維護者通常 1 天內回覆
-
area:tooling bug good first issue priority:P3
難度 2/5 1-3 小時 新手友好度 72/100
michaelnavazhylau/ngspice-rs#129 ·
維護者通常 1 天內回覆