Author eval suite for agent `azure-template-generator`

Open
#94 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
58/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Quiet
Tech stack
yaml
Domain
testing, tooling

Research direction

Read .github/agents/azure-template-generator.agent.md and inspect the eval conventions under .github/evals/ before running /agent-bench azure-template-generator. Author eval.yaml, the positive and negative task files, and the expanded entry in .github/evals/manifest.yaml; run waza run .github/evals/agents/azure-template-generator/eval.yaml -v and confirm mock execution, refusal graders, and a real-model positive result meet the acceptance checklist.

Written by the indexing model from the issue text.

Description

AI-evals enhancement good first issue

Agent

azure-template-generator — source: .github/agents/azure-template-generator.agent.md

Scope

Author the eval suite at .github/evals/agents/azure-template-generator/:

  • eval.yaml — suite config (executor, model, graders)
  • At least 2 positive tasks under tasks/positive-*.yaml
  • At least 1 negative task under tasks/negative-*.yaml
  • Entry added to .github/evals/manifest.yaml at tier: expanded

Notes

Positive tasks that generate ARM templates via create should ALSO require the agent to paste the rendered JSON inline in a fenced code block in its chat response. output_contains graders only check chat output, not files created via create — without an inline echo the grader can't audit the template.

Procedure

  1. /agent-bench azure-template-generator drafts the suite from the live .agent.md.
  2. waza run .github/evals/agents/azure-template-generator/eval.yaml -v locally.
  3. /agent-improve azure-template-generator to iterate on graders.
  4. Open PR.
  5. Mock CI runs automatically. A maintainer will dispatch a real-model run before merge.

Acceptance

  • Suite runs cleanly in mock executor.
  • At least one positive task passes in a real-model run.
  • All negative tasks produce a refusal or out-of-scope acknowledgement.
  • manifest.yaml entry added; PR description includes the real-model run summary.

Conventions to follow

  • Persona lock: refusal graders should accept the agent's own scope language.
  • Prompt graders need continue_session: true in their grader config.

Related

  • Umbrella: #93
  • Harness: #61
Dominant language
JavaScript
Stars
269
Forks
45
Avg merge
1d 4h
Merged PRs (30d)
15

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Azure/git-ape

All issues in Azure/git-ape

Similar issues

More JavaScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.