test(playwright): add a claude plugin eval suite so skill-body changes can re-measure output
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- docker, shell
- Domain
- testing-qa
Research direction
Start with plugins/playbooks/skills/skill-authoring/reference/skill-criteria.md line 191 and the playwright plugin's existing skills/*/evals/evals.json files, plus check-evals-quality.sh, to see what a plugin-eval suite must contain. Design the evals/ suite (prompt.md, case.yaml, graders) with /evals:design, then resolve the Bash and socat/docker host prerequisites. Done means a suite exists that runs on a supported host and a baseline result is recorded against #6650's skill body.
Written by the indexing model from the issue text.
Description
No related PR blocks on this; follow-up from #6650.
Problem
plugins/playbooks/skills/skill-authoring/reference/skill-criteria.md:191 says a skill-body change re-measures output with /evals:plugin-eval. The playwright plugin has no plugin-eval suite (evals/ with prompt.md, case.yaml, graders); it has only skill-creator skills/*/evals/evals.json files. So #6650 (shared login state by default) could only pass the static evals.json check (5 cases, check-evals-quality.sh PASS, 0 warnings), with no measured with-versus-without delta.
Also needed on the run host
The eval sandbox refuses cases that grant Bash on melo WSL hosts: socat is missing (bwrap is present), and ~/.docker holds symbolic links (contexts, features.json). Any Bash-granting suite scores 0 there.
Proposed
- Design the suite with
/evals:design(or generate trigger cases via/skill-quality:check measure-invocationemit-plugin-eval), covering at least the shared-login load-after-open default and the expired-login gotcha. - Resolve the host prerequisites (provisioning) or name a host that can run it.
- Run it once against #6650's skill body as the baseline.
🤖 Generated with Claude Code
- Dominant language
- Shell
- Stars
- 22
- Forks
- 2
- Avg merge
- 5h 15m
- Merged PRs (30d)
- 833
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from melodic-software/claude-code-plugins
-
good first issue needs-triage priority: medium
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
melodic-software/claude-code-plugins#6631 · 1 comment ·
Maintainers usually reply within 1 day
-
needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
melodic-software/claude-code-plugins#6547 ·
Maintainers usually reply within 1 day
-
needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
melodic-software/claude-code-plugins#6535 ·
Maintainers usually reply within 1 day
-
test_comment_census.py: SccArgv flag-shaped-filename test errors on Windows (#!/bin/sh scc shim)Opengood first issue needs-triage priority: low
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
melodic-software/claude-code-plugins#6532 · 1 comment ·
Maintainers usually reply within 1 day
-
good first issue needs-triage priority: low
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
melodic-software/claude-code-plugins#6390 · 1 comment ·
Maintainers usually reply within 1 day
All issues in melodic-software/claude-code-plugins
Similar issues
-
`check_java_version()` fails when Java path contains spaces (Windows / Git Bash, `C:\Program Files`)Open
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
aws-samples/appmod-blueprints#972 ·
Maintainers usually reply within 1 day
-
[Bug]: remote-ls --updates reports up-to-date OCI refs because it ignores deployed Alt-idPossibly taken @Joao-kouznetz claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
status:needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day