Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

ADF judge: bounded decision-model agreement experiment (current judge vs Jev vs Levanto Sage vs local)

未關閉
#965 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

評估

難度
5/5
預估耗時
一週以上
新手友好度
30/100
Issue 類型
功能
描述清晰度
基本清楚
活躍度
活躍
技術堆疊
rust
領域
ai, testing-qa

研究方向

Start with the claude-api build-eval discipline and trace the current ADF judge plus the llm-proxy typesafe::/jev:: adapter. Check how a backend using the /v1/systemone shape could be invoked, then define the sampled cases, permutations, duplicate runs, noise floor, Wilson CIs, and cost-per-task measurements. Done means reporting agreement, cost, permutation swing, and a keep/trial/drop recommendation for every arm.

由索引模型根據 Issue 內容生成。

描述

Goal

A bounded experiment on one ADF judge decision: does a typed decision model agree with the current reasoned verdict, and what does each option cost per completed task (not per call)?

Arms

  1. The current ADF judge verdict (baseline).
  2. TypeSafe Jev Choice (already reachable through the llm-proxy typesafe:: / jev:: adapter).
  3. Levanto Sage choice / yesno, on its free tier (100 decisions a month), with reasoning: "off" and "on" as separate conditions.
  4. A local open-weights backend (openjev-sglang or AnyJev), if one is available on the same /v1/systemone shape.

Method

Build the eval with the claude-api build-eval discipline:

  • Cases sampled from real judge inputs; hard cases chosen by a human, not by current failures.
  • Permute the option order for every decision-model arm. A permutation swing larger than the arm-to-arm difference invalidates that arm's result.
  • Run each grader twice to measure its flip rate.
  • State the noise floor before comparing arms. Report agreement with Wilson CIs, not point estimates.

Out of scope

Replacing the reasoned judge. Neither vendor returns a justification: Sage's meta.reasoning is telemetry only. So a decision-model verdict cannot satisfy the requirement that verdicts be auditable. This experiment measures agreement and cost only.

Acceptance

  • Agreement and cost per task for each arm, with CIs and the permutation swing.
  • A written recommendation: keep, trial, or drop.

Related

  • Lane-policy decision for metered classifiers (separate issue).

Mirror of Gitea terraphim/terraphim-ai#3418 (source of truth).

主要語言
Rust
星號
65
分支
5
平均合併
1 小時 17 分鐘
30 天內合併 PR
2

環境準備

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

terraphim/terraphim-ai 的其他 Issue

查看 terraphim/terraphim-ai 的全部 Issue

相似的 Issue

更多 Rust Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。