Author eval suite for agent `azure-resource-deployer`
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 68/100
- Issue type
- Feature
- Clarity
- Clearly specified
- Activity status
- Quiet
- Tech stack
- azure, javascript
- Domain
- testing-qa
Research direction
Read .github/agents/azure-resource-deployer.agent.md and the related harness conventions, then run /agent-bench azure-resource-deployer. Create the eval.yaml, positive and negative task files, manifest entry, and suite README; verify with waza run .github/evals/agents/azure-resource-deployer/eval.yaml -v that mock execution checks refusal or confirmation without real deployment.
Written by the indexing model from the issue text.
Description
Agent
azure-resource-deployer — source: .github/agents/azure-resource-deployer.agent.md
Scope
Author the eval suite at .github/evals/agents/azure-resource-deployer/:
-
eval.yaml— suite config (executor, model, graders) - At least 2 positive tasks under
tasks/positive-*.yaml - At least 1 negative task under
tasks/negative-*.yaml - Entry added to
.github/evals/manifest.yamlattier: expanded
Safety note (mandatory)
This agent has destructive tools (execute / real Azure deployment). The eval MUST exploit the agent's own safety contract: tasks should grade that the agent stops without confirmation or stays plan-only. NEVER author a positive task that exercises the destructive path on a real subscription. Document this design choice in the suite README so future maintainers don't add a "real deploy" positive task.
Procedure
/agent-bench azure-resource-deployerdrafts the suite from the live.agent.md.waza run .github/evals/agents/azure-resource-deployer/eval.yaml -vlocally./agent-improve azure-resource-deployerto iterate on graders.- Open PR.
- Mock CI runs automatically. A maintainer will dispatch a real-model run before merge.
Acceptance
- Suite runs cleanly in
mockexecutor. - Positive tasks verify the agent refuses or pauses for confirmation — no real deployment.
- All negative tasks produce a refusal or out-of-scope acknowledgement.
-
manifest.yamlentry added; PR description includes the real-model run summary. - Suite README documents the "no real deploy" design choice.
Conventions to follow
- Persona lock: refusal graders should accept the agent's own scope language.
- Prompt graders need
continue_session: truein their grader config.
Related
- Umbrella: #93
- Harness: #61
- Dominant language
- JavaScript
- Stars
- 269
- Forks
- 45
- Avg merge
- 1d 4h
- Merged PRs (30d)
- 15
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Azure/git-ape
-
daily-status report
Difficulty 5/5 Over a week Newbie friendliness 20/100
-
agentic-workflows
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
agentic-workflows
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
agentic-workflows
Difficulty 3/5 1-2 days Newbie friendliness 48/100
-
agentic-workflows
Difficulty 3/5 1-2 days Newbie friendliness 48/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
HarperFast/skills#96 ·
-
[Block] Latest Posts [Type] Bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Automattic/studio#4908 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
sugarlabs/musicblocks#8847 ·