Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

ADF judge: bounded decision-model agreement experiment (current judge vs Jev vs Levanto Sage vs local)

Aberta
#965 2 comentários 0 reações 0 responsáveis Ver no GitHub

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
5/5
Tempo estimado
Mais de uma semana
Facilidade para iniciantes
30/100
Tipo de issue
Funcionalidade
Clareza
Razoavelmente clara
Status de atividade
Ativa
Stack de tecnologia
rust
Domínio
ai, testing-qa

Direção de pesquisa

Start with the claude-api build-eval discipline and trace the current ADF judge plus the llm-proxy typesafe::/jev:: adapter. Check how a backend using the /v1/systemone shape could be invoked, then define the sampled cases, permutations, duplicate runs, noise floor, Wilson CIs, and cost-per-task measurements. Done means reporting agreement, cost, permutation swing, and a keep/trial/drop recommendation for every arm.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

Goal

A bounded experiment on one ADF judge decision: does a typed decision model agree with the current reasoned verdict, and what does each option cost per completed task (not per call)?

Arms

  1. The current ADF judge verdict (baseline).
  2. TypeSafe Jev Choice (already reachable through the llm-proxy typesafe:: / jev:: adapter).
  3. Levanto Sage choice / yesno, on its free tier (100 decisions a month), with reasoning: "off" and "on" as separate conditions.
  4. A local open-weights backend (openjev-sglang or AnyJev), if one is available on the same /v1/systemone shape.

Method

Build the eval with the claude-api build-eval discipline:

  • Cases sampled from real judge inputs; hard cases chosen by a human, not by current failures.
  • Permute the option order for every decision-model arm. A permutation swing larger than the arm-to-arm difference invalidates that arm's result.
  • Run each grader twice to measure its flip rate.
  • State the noise floor before comparing arms. Report agreement with Wilson CIs, not point estimates.

Out of scope

Replacing the reasoned judge. Neither vendor returns a justification: Sage's meta.reasoning is telemetry only. So a decision-model verdict cannot satisfy the requirement that verdicts be auditable. This experiment measures agreement and cost only.

Acceptance

  • Agreement and cost per task for each arm, with CIs and the permutation swing.
  • A written recommendation: keep, trial, or drop.

Related

  • Lane-policy decision for metered classifiers (separate issue).

Mirror of Gitea terraphim/terraphim-ai#3418 (source of truth).

Linguagem predominante
Rust
Estrelas
65
Forks
5
Merge médio
1h 17min
PRs com merge (30d)
2

Preparar o ambiente

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de terraphim/terraphim-ai

Todas as issues de terraphim/terraphim-ai

Issues semelhantes

Mais issues de Rust

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.