Author eval suite for agent `azure-template-generator`
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 58/100
Research direction
Read .github/agents/azure-template-generator.agent.md and inspect the eval conventions under .github/evals/ before running /agent-bench azure-template-generator. Author eval.yaml, the positive and negative task files, and the expanded entry in .github/evals/manifest.yaml; run waza run .github/evals/agents/azure-template-generator/eval.yaml -v and confirm mock execution, refusal graders, and a real-model positive result meet the acceptance checklist.
Written by the indexing model from the issue text.
Description
Agent
azure-template-generator — source: .github/agents/azure-template-generator.agent.md
Scope
Author the eval suite at .github/evals/agents/azure-template-generator/:
-
eval.yaml— suite config (executor, model, graders) - At least 2 positive tasks under
tasks/positive-*.yaml - At least 1 negative task under
tasks/negative-*.yaml - Entry added to
.github/evals/manifest.yamlattier: expanded
Notes
Positive tasks that generate ARM templates via create should ALSO require the agent to paste the rendered JSON inline in a fenced code block in its chat response. output_contains graders only check chat output, not files created via create — without an inline echo the grader can't audit the template.
Procedure
/agent-bench azure-template-generatordrafts the suite from the live.agent.md.waza run .github/evals/agents/azure-template-generator/eval.yaml -vlocally./agent-improve azure-template-generatorto iterate on graders.- Open PR.
- Mock CI runs automatically. A maintainer will dispatch a real-model run before merge.
Acceptance
- Suite runs cleanly in
mockexecutor. - At least one positive task passes in a real-model run.
- All negative tasks produce a refusal or out-of-scope acknowledgement.
-
manifest.yamlentry added; PR description includes the real-model run summary.
Conventions to follow
- Persona lock: refusal graders should accept the agent's own scope language.
- Prompt graders need
continue_session: truein their grader config.
Related
- Umbrella: #93
- Harness: #61
- Dominant language
- JavaScript
- Stars
- 269
- Forks
- 45
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 15
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Azure/git-ape
-
daily-status report
Difficulty 5/5 Over a week Newbie friendliness 20/100
-
agentic-workflows
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
agentic-workflows
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
agentic-workflows
Difficulty 3/5 1-2 days Newbie friendliness 48/100
-
agentic-workflows
Difficulty 3/5 1-2 days Newbie friendliness 48/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
HarperFast/skills#96 ·
-
[Block] Latest Posts [Type] Bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Automattic/studio#4908 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
sugarlabs/musicblocks#8847 ·