Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

fix(skill-evals): 32 eval steps without a user-prompt-template.md get security-issue-import's prompt

Open
#1,492 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

@dpol1 is already working on this.

Since Oct 5, 2026.

  • #1524 by @dpol1 — open

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
72/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
testing, tooling

Research direction

Start in tools/skill-evals/src/skill_evals/runner.py:194-197 and inspect the USER_PROMPT_TEMPLATE definition at lines 96-110, then compare the affected eval directories and security-issue-import/step-2a-semantic-sweep. Give the security step its explicit template, make the fallback neutral, and re-run the 32 affected steps; done means the prompts are appropriate and any verdict changes are recorded.

Written by the indexing model from the issue text.

Description

What

When a fixtures directory has no user-prompt-template.md, the runner falls back to USER_PROMPT_TEMPLATE (tools/skill-evals/src/skill_evals/runner.py:194-197). That default (lines 96-110) is the security-issue-import template:

## Existing open trackers (corpus)
{corpus}
## Reporter roster (existing trackers mapped to reporter email)
{roster}
## Incoming report
{report}
Apply the semantic sweep and reporter-identity check. Return JSON only.

33 step directories under tools/skill-evals/evals/ have no template. One is security-issue-import/step-2a-semantic-sweep, the step the default was written for. The other 32 belong to other skills:

  • every step of release-announce-draft, release-archive-sweep, release-audit-report, release-keys-sync, release-prepare, release-rc-cut, release-verify-rc and release-vote-tally (30 steps);
  • reviewer-routing/step-0-preflight;
  • non-asf-profile-smoke/step-release-backend-preflight.

For those steps the model gets an empty "Existing open trackers" section, an empty "Reporter roster" section, the case input under "Incoming report", and an instruction to run a semantic sweep and a reporter-identity check that has nothing to do with the step under test.

Why it matters

The system prompt is correct: it carries the step extracted from the skill. The user prompt is not: it asks for an import-skill task. A PASS on these steps says less than it appears to, and a FAIL may come from the prompt rather than the skill. AGENTS.md makes the eval suite the check for every skill change, and almost all of the affected steps are in the release family. This does not mean these evals fail; I have not measured how many verdicts change with a neutral prompt.

Suggested fix
  1. Make the default template neutral: the {report} slot and "Return JSON only.", nothing else. str.format() ignores the unused corpus / roster arguments, so the call site does not change.
  2. Give security-issue-import/step-2a-semantic-sweep its own user-prompt-template.md holding today's default, so its prompt stays the same.
  3. Re-run the 32 affected steps and note any verdict that moves.

The alternative is to make the template mandatory and have the runner fail when it is missing, which means 32 new files.

Found via

Reading the runner's prompt assembly while mapping the eval harness. The 33 directories are those with a step-config.json and no user-prompt-template.md on main (37670575).

Dominant language
Python
Stars
108
Forks
94
Avg merge
7h 21m
Merged PRs (30d)
296

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/magpie

All issues in apache/magpie

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.