Hacktoberfest 2026 : les issues que les mainteneurs ont marquées pour octobre, ouvertes et accessibles aux débutants. Parcourir les issues Hacktoberfest

ADF judge: bounded decision-model agreement experiment (current judge vs Jev vs Levanto Sage vs local)

Ouverte
#965 2 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Personne n'a encore pris cette issue.

Évaluation

Difficulté
5/5
Temps estimé
Plus d'une semaine
Accessibilité débutants
30/100
Type d'issue
Fonctionnalité
Clarté
Plutôt claire
Activité
Active
Stack technique
rust
Domaine
ai, testing-qa

Piste de recherche

Start with the claude-api build-eval discipline and trace the current ADF judge plus the llm-proxy typesafe::/jev:: adapter. Check how a backend using the /v1/systemone shape could be invoked, then define the sampled cases, permutations, duplicate runs, noise floor, Wilson CIs, and cost-per-task measurements. Done means reporting agreement, cost, permutation swing, and a keep/trial/drop recommendation for every arm.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Description

Goal

A bounded experiment on one ADF judge decision: does a typed decision model agree with the current reasoned verdict, and what does each option cost per completed task (not per call)?

Arms

  1. The current ADF judge verdict (baseline).
  2. TypeSafe Jev Choice (already reachable through the llm-proxy typesafe:: / jev:: adapter).
  3. Levanto Sage choice / yesno, on its free tier (100 decisions a month), with reasoning: "off" and "on" as separate conditions.
  4. A local open-weights backend (openjev-sglang or AnyJev), if one is available on the same /v1/systemone shape.

Method

Build the eval with the claude-api build-eval discipline:

  • Cases sampled from real judge inputs; hard cases chosen by a human, not by current failures.
  • Permute the option order for every decision-model arm. A permutation swing larger than the arm-to-arm difference invalidates that arm's result.
  • Run each grader twice to measure its flip rate.
  • State the noise floor before comparing arms. Report agreement with Wilson CIs, not point estimates.

Out of scope

Replacing the reasoned judge. Neither vendor returns a justification: Sage's meta.reasoning is telemetry only. So a decision-model verdict cannot satisfy the requirement that verdicts be auditable. This experiment measures agreement and cost only.

Acceptance

  • Agreement and cost per task for each arm, with CIs and the permutation swing.
  • A written recommendation: keep, trial, or drop.

Related

  • Lane-policy decision for metered classifiers (separate issue).

Mirror of Gitea terraphim/terraphim-ai#3418 (source of truth).

Langage dominant
Rust
Étoiles
65
Forks
5
Merge moyen
1 h 17 min
PR mergées (30 j)
2

Préparer son environnement

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de terraphim/terraphim-ai

Toutes les issues de terraphim/terraphim-ai

Issues similaires

Plus d'issues Rust

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.