fix(skill-evals): 32 eval steps without a user-prompt-template.md get security-issue-import's prompt
Maintainers usually reply within 1 day
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 72/100
Research direction
Start in tools/skill-evals/src/skill_evals/runner.py:194-197 and inspect the USER_PROMPT_TEMPLATE definition at lines 96-110, then compare the affected eval directories and security-issue-import/step-2a-semantic-sweep. Give the security step its explicit template, make the fallback neutral, and re-run the 32 affected steps; done means the prompts are appropriate and any verdict changes are recorded.
Written by the indexing model from the issue text.
Description
What
When a fixtures directory has no user-prompt-template.md, the runner falls back to USER_PROMPT_TEMPLATE (tools/skill-evals/src/skill_evals/runner.py:194-197). That default (lines 96-110) is the security-issue-import template:
## Existing open trackers (corpus)
{corpus}
## Reporter roster (existing trackers mapped to reporter email)
{roster}
## Incoming report
{report}
Apply the semantic sweep and reporter-identity check. Return JSON only.
33 step directories under tools/skill-evals/evals/ have no template. One is security-issue-import/step-2a-semantic-sweep, the step the default was written for. The other 32 belong to other skills:
- every step of
release-announce-draft,release-archive-sweep,release-audit-report,release-keys-sync,release-prepare,release-rc-cut,release-verify-rcandrelease-vote-tally(30 steps); reviewer-routing/step-0-preflight;non-asf-profile-smoke/step-release-backend-preflight.
For those steps the model gets an empty "Existing open trackers" section, an empty "Reporter roster" section, the case input under "Incoming report", and an instruction to run a semantic sweep and a reporter-identity check that has nothing to do with the step under test.
Why it matters
The system prompt is correct: it carries the step extracted from the skill. The user prompt is not: it asks for an import-skill task. A PASS on these steps says less than it appears to, and a FAIL may come from the prompt rather than the skill. AGENTS.md makes the eval suite the check for every skill change, and almost all of the affected steps are in the release family. This does not mean these evals fail; I have not measured how many verdicts change with a neutral prompt.
Suggested fix
- Make the default template neutral: the
{report}slot and "Return JSON only.", nothing else.str.format()ignores the unusedcorpus/rosterarguments, so the call site does not change. - Give
security-issue-import/step-2a-semantic-sweepits ownuser-prompt-template.mdholding today's default, so its prompt stays the same. - Re-run the 32 affected steps and note any verdict that moves.
The alternative is to make the template mandatory and have the runner fail when it is missing, which means 32 new files.
Found via
Reading the runner's prompt assembly while mapping the eval harness. The 33 directories are those with a step-config.json and no user-prompt-template.md on main (37670575).
- Dominant language
- Python
- Stars
- 108
- Forks
- 94
- Avg merge
- 7h 21m
- Merged PRs (30d)
- 296
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/magpie
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-2 days Newbie friendliness 64/100
Maintainers usually reply within 1 day
-
capability:platform enhancement family:issue family:pairing family:pr-management family:release-management family:security family:setup family:utilities
Difficulty 5/5 Over a week Newbie friendliness 35/100
Maintainers usually reply within 1 day
-
fix(bitbucket): Cloud pr status omits the pull request state and head commitPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 3/5 1-2 days Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
fix(pr-triage): the mark-ready guard can be bypassed with a repeated or comma-separated --add-labelPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 3/5 1-2 days Newbie friendliness 76/100
Maintainers usually reply within 1 day
Similar issues
-
bug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
debpalash/VoiceStudio#2624 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Make Catch2 optional when `RDK_BUILD_CPP_TESTS=OFF`Possibly taken @pechersky claimed this today. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 2 days
-
There are a few redundant calls to `fdesc._setCloseOnExec()`Possibly taken @gudnimg claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day