feat(evals): run plugin eval suites in a CI lane and hill-climb in the background
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 28/100
Línea de trabajo
Start from .github/workflows/ (especially pr-require-checks.yml) and the plugin-eval CI notes in /evals:plugin-eval plus reference/ci.md (exit-code parser, partial/skippedPaidGraders). Read ADR 0049 and the plugin-evals CI docs for --trust-plugin, --json, --max-cost-usd, and sandbox (bubblewrap/socat). Done for the first slice is an advisory PR job that validates with validate-cases.py, runs only deterministic graders for the changed plugin, and posts a with-versus-without table plus skill-load/firing health—not auto-fix or hill-climb.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Problem
Every claude plugin eval run in this repository is started by hand in the operator's terminal and paid for there. No workflow under .github/workflows/ runs a plugin eval suite; pr-require-checks.yml only lints skill-creator evals/evals.json files and fixtures. Sessions cannot run the suites either: a worktree-isolated session refuses any command containing eval (#5696, #6374), and a background run dies with its launcher (#6333). So a change to a skill that has a suite (performance, improvement, source-control, code-tidying, education, discovery, planning, evals) merges without its with-versus-without delta unless someone remembers to run it.
Evidence from one session (2026-10-04)
These are the failure modes an automated lane has to handle, not only the reason to build one.
- All runs were manual. The session's worktree-isolation guard refused
claude plugin eval, so every run was pasted into the user's terminal and paid for there. - Several suites at first measured nothing.
- The skill never loaded in some cases:
Bashwas denied during the fanout skill's shell preprocessing, so the with-arm ran without the skill body. - The skill-fired grader misread slash-invoked runs.
- Firing varied widely between runs: implement-dispatch fired 6/12, then 0/3, then 6/6.
- The skill never loaded in some cases:
- Judge graders needed calibration. One judge scored 89%, then 80% agreement, and was replaced by regex graders.
- Confirming runs found real effects. Four changes each moved a case from 0/3 to 3/3.
A lane that reports a delta without checking that the skill loaded, that the grader reads the run shape correctly, and that the run count is enough for the observed variance will post confident numbers that mean nothing.
Can it run in Actions today
Yes, as a plain workflow step. https://code.claude.com/docs/en/plugin-evals#run-evals-in-ci (fetched 2026-10-04) documents a CI invocation with --trust-plugin, --json, --threshold, pinned --model and --judge-model, --no-publish and --max-cost-usd, and says a runner needs a Claude Code install and credentials in the environment. Each case runs as its own non-interactive session with only the granted tools, so no --permission-mode applies to the eval step itself. Cases that grant Bash, Write or Edit need the OS sandbox: on Linux, bubblewrap and socat (Grant tools). Whether those run on a GitHub-hosted Ubuntu runner (unprivileged user namespaces) is not verified.
The auto-fix and hill-climb parts would be model steps, which https://code.claude.com/docs/en/github-actions runs through claude-code-action with a prompt and claude_args. Those fall under ADR 0049.
Our runner notes: /evals:plugin-eval CI section and its reference/ci.md (exit-code parser, partial and skippedPaidGraders handling, the "currently unavailable" server-side switch). For spend, see https://code.claude.com/docs/en/costs and https://platform.claude.com/docs/en/about-claude/pricing; this issue states no figures.
Proposed lane
- Trigger:
pull_requeston same-repository branches that touchplugins/<p>/skills/**,plugins/<p>/agents/**orplugins/<p>/evals/**, for plugins with a suite. Skip drafts, matching the other lanes. Aworkflow_dispatchinput for a full sweep, and later a schedule for judge-graded cases. - Scope: per changed plugin, only that plugin's suite. No cross-plugin runs on a PR.
- Authority it must not have:
- No merge, no
ci-statuswrite, nochecks: writeorworkflowspermission in any model step (ADR 0049 rule 3). - Any fix, new case, or skill edit is pushed as a draft PR or a commit to a lane branch, never to the PR's branch without the author's opt-in and never to
main. - PR titles, bodies, comments and eval transcripts reaching a model step are framed as data per
docs/conventions/untrusted-content/README.md; trusted-actor and same-repo gates per ADR 0049 rules 1 and 2; the org kill switch per rule 4.
- No merge, no
Smallest first slice
On a PR that touches plugins/<p>/skills, agents or evals:
- Run
/evals:validate(validate-cases.py) on that plugin's suite. No model call. - Run that plugin's suite with deterministic graders only,
--max-cost-usdset, both models pinned,--json. - Post the with-versus-without table as an advisory check or job summary, gated through the
reference/ci.mdparser so apartial,skippedPaidGraders, errored or delta-less case reads NOT COMPARABLE instead of a score. - Add one health line per case: did the skill load in the with-arm (from the trace), and how many runs fired.
No auto-fix, no new cases, no hill-climb in this slice. Those come after the advisory numbers have been read on real PRs for a while.
Later slices (not in scope yet)
- Author or repair cases for skills whose suite measured nothing (skill not loaded, grader misreading the run shape), as draft PRs.
- Judge calibration: run a judge against its must-pass and must-fail samples and refuse to use it below a set agreement.
- Background hill-climb on a schedule: propose skill edits, keep only those that beat the baseline over enough runs, open each as a draft PR.
Open questions
- Sandbox on hosted runners. Do
bubblewrapandsocatwork onubuntu-latest, or do Bash-granting cases need a container or--ablation noneread-only suites only? Settled by one dispatch run of a Bash-granting case. - Spend control. Which credential (API key vs subscription token, per the github-actions page), what
--max-cost-usdper PR, and whether a label opts a PR in rather than running on every push. Needs the operator's decision. - Flakes. With firing at 6/12 then 0/3 on the same skill, what run count makes a PR delta trustworthy, and does the lane report an interval instead of a point?
/evals:plugin-evalreading-results guidance is the starting point. - Judge calibration gating. Should judge-graded cases stay out of the PR lane until each judge passes its calibration samples (#5879 lists known gaps)?
- Server-side switch.
plugin eval is currently unavailablecan fire with no local fix, so the lane must stay advisory and never become a required check.
Related
- #5696, #6374, #6333: why sessions cannot run suites today.
- #5879: judge calibration gaps and converting other plugins' evals.
- ADR 0049 (
docs/adr/0049-run-ci-lanes-on-github-hosted-runners-under-trigger-and-token-hardening.md): hosted-runner lane hardening.
- Lenguaje dominante
- Shell
- Estrellas
- 22
- Forks
- 2
- Merge medio
- 5 h 11 min
- PR fusionados (30 d)
- 838
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de melodic-software/claude-code-plugins
-
good first issue needs-triage priority: medium
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
melodic-software/claude-code-plugins#6631 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
needs-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
melodic-software/claude-code-plugins#6547 ·
Los mantenedores suelen responder en 1 día
-
needs-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
melodic-software/claude-code-plugins#6535 ·
Los mantenedores suelen responder en 1 día
-
test_comment_census.py: SccArgv flag-shaped-filename test errors on Windows (#!/bin/sh scc shim)Abiertogood first issue needs-triage priority: low
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
melodic-software/claude-code-plugins#6532 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
good first issue needs-triage priority: low
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
melodic-software/claude-code-plugins#6390 · 1 comentario ·
Los mantenedores suelen responder en 1 día
Todos los issues de melodic-software/claude-code-plugins
Issues similares
-
Lid close does not lock the session on Apple Silicon (lid-close bind skips omarchy-system-lid-close)Abierto
Dificultad 1/5 1-3 horas Aptitud para principiantes 90/100
omacom/omarchy-mac#701 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
A 20.x release after 21.0.0 would move `latest` back to 20.x, and `next` stays on the release candidatePosiblemente ocupada @armando-navarro la tomó hoy. Abiertocomp: build/pipeline type: bug version: current (v17+)
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
angular/angularfire#3790 ·
Los mantenedores suelen responder en 3 días
-
ready-for-agent
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
LucasSantana-Dev/Lucky#2698 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
collabnix/awesome-mcp-lists#179 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100