ADF judge: bounded decision-model agreement experiment (current judge vs Jev vs Levanto Sage vs local)
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 30/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- rust
- Ambito
- ai, testing-qa
Direzione di ricerca
Start with the claude-api build-eval discipline and trace the current ADF judge plus the llm-proxy typesafe::/jev:: adapter. Check how a backend using the /v1/systemone shape could be invoked, then define the sampled cases, permutations, duplicate runs, noise floor, Wilson CIs, and cost-per-task measurements. Done means reporting agreement, cost, permutation swing, and a keep/trial/drop recommendation for every arm.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Goal
A bounded experiment on one ADF judge decision: does a typed decision model agree with the current reasoned verdict, and what does each option cost per completed task (not per call)?
Arms
- The current ADF judge verdict (baseline).
- TypeSafe Jev
Choice(already reachable through the llm-proxytypesafe::/jev::adapter). - Levanto Sage
choice/yesno, on its free tier (100 decisions a month), withreasoning: "off"and"on"as separate conditions. - A local open-weights backend (openjev-sglang or AnyJev), if one is available on the same
/v1/systemoneshape.
Method
Build the eval with the claude-api build-eval discipline:
- Cases sampled from real judge inputs; hard cases chosen by a human, not by current failures.
- Permute the option order for every decision-model arm. A permutation swing larger than the arm-to-arm difference invalidates that arm's result.
- Run each grader twice to measure its flip rate.
- State the noise floor before comparing arms. Report agreement with Wilson CIs, not point estimates.
Out of scope
Replacing the reasoned judge. Neither vendor returns a justification: Sage's meta.reasoning is telemetry only. So a decision-model verdict cannot satisfy the requirement that verdicts be auditable. This experiment measures agreement and cost only.
Acceptance
- Agreement and cost per task for each arm, with CIs and the permutation swing.
- A written recommendation: keep, trial, or drop.
Related
- Lane-policy decision for metered classifiers (separate issue).
Mirror of Gitea terraphim/terraphim-ai#3418 (source of truth).
- Lingua principale
- Rust
- Stelle
- 65
- Fork
- 5
- Merge medio
- 1h 17m
- PR unite (30g)
- 2
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di terraphim/terraphim-ai
-
terraphim_rlm: require a per-sub-result verification plan (recheck / control / independent route)Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
terraphim/terraphim-ai#969 · 1 commento ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
terraphim/terraphim-ai#968 · 1 commento ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
terraphim/terraphim-ai#967 ·
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
terraphim/terraphim-ai#966 ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
terraphim/terraphim-ai#885 ·
Tutte le issue di terraphim/terraphim-ai
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
I maintainer di solito rispondono entro 5 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
tauri-apps/tauri#16219 ·
I maintainer di solito rispondono entro 2 giorni
-
state:triage-needed
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 1 giorno
-
ktuner keeps a stale ledger path and can never restore that entryForse già presa @Frun1na l’ha presa oggi. Apertacomponent:ktuner
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
agentic-os-org/ANOLISA#6483 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno