Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

feat(evals): run plugin eval suites in a CI lane and hill-climb in the background

Abierto
#6,405 1 comentario 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
28/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
github-actions, python, shell
Área
ci-cd, devops, testing

Línea de trabajo

Start from .github/workflows/ (especially pr-require-checks.yml) and the plugin-eval CI notes in /evals:plugin-eval plus reference/ci.md (exit-code parser, partial/skippedPaidGraders). Read ADR 0049 and the plugin-evals CI docs for --trust-plugin, --json, --max-cost-usd, and sandbox (bubblewrap/socat). Done for the first slice is an advisory PR job that validates with validate-cases.py, runs only deterministic graders for the changed plugin, and posts a with-versus-without table plus skill-load/firing health—not auto-fix or hill-climb.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

needs-human needs-triage

Problem

Every claude plugin eval run in this repository is started by hand in the operator's terminal and paid for there. No workflow under .github/workflows/ runs a plugin eval suite; pr-require-checks.yml only lints skill-creator evals/evals.json files and fixtures. Sessions cannot run the suites either: a worktree-isolated session refuses any command containing eval (#5696, #6374), and a background run dies with its launcher (#6333). So a change to a skill that has a suite (performance, improvement, source-control, code-tidying, education, discovery, planning, evals) merges without its with-versus-without delta unless someone remembers to run it.

Evidence from one session (2026-10-04)

These are the failure modes an automated lane has to handle, not only the reason to build one.

  • All runs were manual. The session's worktree-isolation guard refused claude plugin eval, so every run was pasted into the user's terminal and paid for there.
  • Several suites at first measured nothing.
    • The skill never loaded in some cases: Bash was denied during the fanout skill's shell preprocessing, so the with-arm ran without the skill body.
    • The skill-fired grader misread slash-invoked runs.
    • Firing varied widely between runs: implement-dispatch fired 6/12, then 0/3, then 6/6.
  • Judge graders needed calibration. One judge scored 89%, then 80% agreement, and was replaced by regex graders.
  • Confirming runs found real effects. Four changes each moved a case from 0/3 to 3/3.

A lane that reports a delta without checking that the skill loaded, that the grader reads the run shape correctly, and that the run count is enough for the observed variance will post confident numbers that mean nothing.

Can it run in Actions today

Yes, as a plain workflow step. https://code.claude.com/docs/en/plugin-evals#run-evals-in-ci (fetched 2026-10-04) documents a CI invocation with --trust-plugin, --json, --threshold, pinned --model and --judge-model, --no-publish and --max-cost-usd, and says a runner needs a Claude Code install and credentials in the environment. Each case runs as its own non-interactive session with only the granted tools, so no --permission-mode applies to the eval step itself. Cases that grant Bash, Write or Edit need the OS sandbox: on Linux, bubblewrap and socat (Grant tools). Whether those run on a GitHub-hosted Ubuntu runner (unprivileged user namespaces) is not verified.

The auto-fix and hill-climb parts would be model steps, which https://code.claude.com/docs/en/github-actions runs through claude-code-action with a prompt and claude_args. Those fall under ADR 0049.

Our runner notes: /evals:plugin-eval CI section and its reference/ci.md (exit-code parser, partial and skippedPaidGraders handling, the "currently unavailable" server-side switch). For spend, see https://code.claude.com/docs/en/costs and https://platform.claude.com/docs/en/about-claude/pricing; this issue states no figures.

Proposed lane

  • Trigger: pull_request on same-repository branches that touch plugins/<p>/skills/**, plugins/<p>/agents/** or plugins/<p>/evals/**, for plugins with a suite. Skip drafts, matching the other lanes. A workflow_dispatch input for a full sweep, and later a schedule for judge-graded cases.
  • Scope: per changed plugin, only that plugin's suite. No cross-plugin runs on a PR.
  • Authority it must not have:
    • No merge, no ci-status write, no checks: write or workflows permission in any model step (ADR 0049 rule 3).
    • Any fix, new case, or skill edit is pushed as a draft PR or a commit to a lane branch, never to the PR's branch without the author's opt-in and never to main.
    • PR titles, bodies, comments and eval transcripts reaching a model step are framed as data per docs/conventions/untrusted-content/README.md; trusted-actor and same-repo gates per ADR 0049 rules 1 and 2; the org kill switch per rule 4.

Smallest first slice

On a PR that touches plugins/<p>/skills, agents or evals:

  1. Run /evals:validate (validate-cases.py) on that plugin's suite. No model call.
  2. Run that plugin's suite with deterministic graders only, --max-cost-usd set, both models pinned, --json.
  3. Post the with-versus-without table as an advisory check or job summary, gated through the reference/ci.md parser so a partial, skippedPaidGraders, errored or delta-less case reads NOT COMPARABLE instead of a score.
  4. Add one health line per case: did the skill load in the with-arm (from the trace), and how many runs fired.

No auto-fix, no new cases, no hill-climb in this slice. Those come after the advisory numbers have been read on real PRs for a while.

Later slices (not in scope yet)

  • Author or repair cases for skills whose suite measured nothing (skill not loaded, grader misreading the run shape), as draft PRs.
  • Judge calibration: run a judge against its must-pass and must-fail samples and refuse to use it below a set agreement.
  • Background hill-climb on a schedule: propose skill edits, keep only those that beat the baseline over enough runs, open each as a draft PR.

Open questions

  1. Sandbox on hosted runners. Do bubblewrap and socat work on ubuntu-latest, or do Bash-granting cases need a container or --ablation none read-only suites only? Settled by one dispatch run of a Bash-granting case.
  2. Spend control. Which credential (API key vs subscription token, per the github-actions page), what --max-cost-usd per PR, and whether a label opts a PR in rather than running on every push. Needs the operator's decision.
  3. Flakes. With firing at 6/12 then 0/3 on the same skill, what run count makes a PR delta trustworthy, and does the lane report an interval instead of a point? /evals:plugin-eval reading-results guidance is the starting point.
  4. Judge calibration gating. Should judge-graded cases stay out of the PR lane until each judge passes its calibration samples (#5879 lists known gaps)?
  5. Server-side switch. plugin eval is currently unavailable can fire with no local fix, so the lane must stay advisory and never become a required check.

Related

  • #5696, #6374, #6333: why sessions cannot run suites today.
  • #5879: judge calibration gaps and converting other plugins' evals.
  • ADR 0049 (docs/adr/0049-run-ci-lanes-on-github-hosted-runners-under-trigger-and-token-hardening.md): hosted-runner lane hardening.
Lenguaje dominante
Shell
Estrellas
22
Forks
2
Merge medio
5 h 11 min
PR fusionados (30 d)
838

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de melodic-software/claude-code-plugins

Todos los issues de melodic-software/claude-code-plugins

Issues similares

Más issues de Shell/Bash

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.