Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

ADF judge: bounded decision-model agreement experiment (current judge vs Jev vs Levanto Sage vs local)

Abierto
#965 2 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
30/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
rust
Área
ai, testing-qa

Línea de trabajo

Start with the claude-api build-eval discipline and trace the current ADF judge plus the llm-proxy typesafe::/jev:: adapter. Check how a backend using the /v1/systemone shape could be invoked, then define the sampled cases, permutations, duplicate runs, noise floor, Wilson CIs, and cost-per-task measurements. Done means reporting agreement, cost, permutation swing, and a keep/trial/drop recommendation for every arm.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Goal

A bounded experiment on one ADF judge decision: does a typed decision model agree with the current reasoned verdict, and what does each option cost per completed task (not per call)?

Arms

  1. The current ADF judge verdict (baseline).
  2. TypeSafe Jev Choice (already reachable through the llm-proxy typesafe:: / jev:: adapter).
  3. Levanto Sage choice / yesno, on its free tier (100 decisions a month), with reasoning: "off" and "on" as separate conditions.
  4. A local open-weights backend (openjev-sglang or AnyJev), if one is available on the same /v1/systemone shape.

Method

Build the eval with the claude-api build-eval discipline:

  • Cases sampled from real judge inputs; hard cases chosen by a human, not by current failures.
  • Permute the option order for every decision-model arm. A permutation swing larger than the arm-to-arm difference invalidates that arm's result.
  • Run each grader twice to measure its flip rate.
  • State the noise floor before comparing arms. Report agreement with Wilson CIs, not point estimates.

Out of scope

Replacing the reasoned judge. Neither vendor returns a justification: Sage's meta.reasoning is telemetry only. So a decision-model verdict cannot satisfy the requirement that verdicts be auditable. This experiment measures agreement and cost only.

Acceptance

  • Agreement and cost per task for each arm, with CIs and the permutation swing.
  • A written recommendation: keep, trial, or drop.

Related

  • Lane-policy decision for metered classifiers (separate issue).

Mirror of Gitea terraphim/terraphim-ai#3418 (source of truth).

Lenguaje dominante
Rust
Estrellas
65
Forks
5
Merge medio
1 h 17 min
PR fusionados (30 d)
2

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de terraphim/terraphim-ai

Todos los issues de terraphim/terraphim-ai

Issues similares

Más issues de Rust

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.