ADF judge: bounded decision-model agreement experiment (current judge vs Jev vs Levanto Sage vs local)
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 5/5
- Geschätzter Aufwand
- Über eine Woche
- Anfängerfreundlichkeit
- 30/100
- Issue-Typ
- Feature
- Klarheit
- Größtenteils klar
- Aktivitätsstatus
- Aktiv
- Tech-Stack
- rust
- Bereich
- ai, testing-qa
Rechercherichtung
Start with the claude-api build-eval discipline and trace the current ADF judge plus the llm-proxy typesafe::/jev:: adapter. Check how a backend using the /v1/systemone shape could be invoked, then define the sampled cases, permutations, duplicate runs, noise floor, Wilson CIs, and cost-per-task measurements. Done means reporting agreement, cost, permutation swing, and a keep/trial/drop recommendation for every arm.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Goal
A bounded experiment on one ADF judge decision: does a typed decision model agree with the current reasoned verdict, and what does each option cost per completed task (not per call)?
Arms
- The current ADF judge verdict (baseline).
- TypeSafe Jev
Choice(already reachable through the llm-proxytypesafe::/jev::adapter). - Levanto Sage
choice/yesno, on its free tier (100 decisions a month), withreasoning: "off"and"on"as separate conditions. - A local open-weights backend (openjev-sglang or AnyJev), if one is available on the same
/v1/systemoneshape.
Method
Build the eval with the claude-api build-eval discipline:
- Cases sampled from real judge inputs; hard cases chosen by a human, not by current failures.
- Permute the option order for every decision-model arm. A permutation swing larger than the arm-to-arm difference invalidates that arm's result.
- Run each grader twice to measure its flip rate.
- State the noise floor before comparing arms. Report agreement with Wilson CIs, not point estimates.
Out of scope
Replacing the reasoned judge. Neither vendor returns a justification: Sage's meta.reasoning is telemetry only. So a decision-model verdict cannot satisfy the requirement that verdicts be auditable. This experiment measures agreement and cost only.
Acceptance
- Agreement and cost per task for each arm, with CIs and the permutation swing.
- A written recommendation: keep, trial, or drop.
Related
- Lane-policy decision for metered classifiers (separate issue).
Mirror of Gitea terraphim/terraphim-ai#3418 (source of truth).
- Vorherrschende Sprache
- Rust
- Sterne
- 65
- Forks
- 5
- Ø Merge
- 1 Std. 17 Min.
- Gemergte PRs (30 T.)
- 2
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus terraphim/terraphim-ai
-
terraphim_rlm: require a per-sub-result verification plan (recheck / control / independent route)Offen
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 48/100
terraphim/terraphim-ai#969 · 1 Kommentar ·
-
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
terraphim/terraphim-ai#968 · 1 Kommentar ·
-
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 35/100
terraphim/terraphim-ai#967 ·
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 48/100
terraphim/terraphim-ai#966 ·
-
Schwierigkeit 5/5 Über eine Woche Anfängerfreundlichkeit 25/100
terraphim/terraphim-ai#885 ·
Alle Issues in terraphim/terraphim-ai
Ähnliche Issues
-
app enhancement
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 80/100
elodin-sys/elodin#890 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 85/100
guidance-ai/llguidance#391 ·
-
documentation
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 85/100
Verifiedz/Shimmer#144 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
点击设置提示`操作未能完成,详情请查看应用日志`Offen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
Y-ASLant/ElegantClipboard#166 ·