Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

ADF judge: bounded decision-model agreement experiment (current judge vs Jev vs Levanto Sage vs local)

Offen
#965 2 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
5/5
Geschätzter Aufwand
Über eine Woche
Anfängerfreundlichkeit
30/100
Issue-Typ
Feature
Klarheit
Größtenteils klar
Aktivitätsstatus
Aktiv
Tech-Stack
rust
Bereich
ai, testing-qa

Rechercherichtung

Start with the claude-api build-eval discipline and trace the current ADF judge plus the llm-proxy typesafe::/jev:: adapter. Check how a backend using the /v1/systemone shape could be invoked, then define the sampled cases, permutations, duplicate runs, noise floor, Wilson CIs, and cost-per-task measurements. Done means reporting agreement, cost, permutation swing, and a keep/trial/drop recommendation for every arm.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Goal

A bounded experiment on one ADF judge decision: does a typed decision model agree with the current reasoned verdict, and what does each option cost per completed task (not per call)?

Arms

  1. The current ADF judge verdict (baseline).
  2. TypeSafe Jev Choice (already reachable through the llm-proxy typesafe:: / jev:: adapter).
  3. Levanto Sage choice / yesno, on its free tier (100 decisions a month), with reasoning: "off" and "on" as separate conditions.
  4. A local open-weights backend (openjev-sglang or AnyJev), if one is available on the same /v1/systemone shape.

Method

Build the eval with the claude-api build-eval discipline:

  • Cases sampled from real judge inputs; hard cases chosen by a human, not by current failures.
  • Permute the option order for every decision-model arm. A permutation swing larger than the arm-to-arm difference invalidates that arm's result.
  • Run each grader twice to measure its flip rate.
  • State the noise floor before comparing arms. Report agreement with Wilson CIs, not point estimates.

Out of scope

Replacing the reasoned judge. Neither vendor returns a justification: Sage's meta.reasoning is telemetry only. So a decision-model verdict cannot satisfy the requirement that verdicts be auditable. This experiment measures agreement and cost only.

Acceptance

  • Agreement and cost per task for each arm, with CIs and the permutation swing.
  • A written recommendation: keep, trial, or drop.

Related

  • Lane-policy decision for metered classifiers (separate issue).

Mirror of Gitea terraphim/terraphim-ai#3418 (source of truth).

Vorherrschende Sprache
Rust
Sterne
65
Forks
5
Ø Merge
1 Std. 17 Min.
Gemergte PRs (30 T.)
2

Entwicklungsumgebung

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus terraphim/terraphim-ai

Alle Issues in terraphim/terraphim-ai

Ähnliche Issues

Weitere Issues zu Rust

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.