Author eval suite for agent `git-ape`

Open
#107 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
58/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Quiet
Tech stack
github, yaml

Research direction

Read .github/agents/git-ape.agent.md and compare the established sub-agent suites once their baselines are stable. Run /agent-bench git-ape, then validate .github/evals/agents/git-ape/eval.yaml with waza run ... -v; done means the mock suite is clean, real-model and negative-task acceptance checks pass, and .github/evals/manifest.yaml includes the expanded-tier entry.

Written by the indexing model from the issue text.

Description

AI-evals enhancement

Agent

git-ape — source: .github/agents/git-ape.agent.md

Scope

Author the eval suite at .github/evals/agents/git-ape/:

  • eval.yaml — suite config (executor, model, graders)
  • At least 2 positive tasks under tasks/positive-*.yaml
  • At least 1 negative task under tasks/negative-*.yaml
  • Entry added to .github/evals/manifest.yaml at tier: expanded

Dependency note

git-ape is the orchestrator agent. Defer this suite until most sub-agent suites (azure-requirements-gatherer, azure-template-generator, azure-resource-deployer, azure-iac-exporter) are stable — regressions in this suite are easier to root-cause when each sub-agent has its own established baseline.

Procedure

  1. /agent-bench git-ape drafts the suite from the live .agent.md.
  2. waza run .github/evals/agents/git-ape/eval.yaml -v locally.
  3. /agent-improve git-ape to iterate on graders.
  4. Open PR.
  5. Mock CI runs automatically. A maintainer will dispatch a real-model run before merge.

Acceptance

  • Suite runs cleanly in mock executor.
  • At least one positive task passes in a real-model run.
  • All negative tasks produce a refusal or out-of-scope acknowledgement.
  • manifest.yaml entry added; PR description includes the real-model run summary.

Conventions to follow

  • Persona lock: refusal graders should accept the agent's own scope language.
  • Prompt graders need continue_session: true in their grader config.
  • This agent has destructive tools through delegation. Apply the same "no real deploy" rule as azure-resource-deployer: positive tasks grade safety-contract behavior, not real Azure execution.

Related

  • Umbrella: #93
  • Harness: #61
Dominant language
JavaScript
Stars
269
Forks
45
Avg merge
1d 4h
Merged PRs (30d)
15

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Azure/git-ape

All issues in Azure/git-ape

Similar issues

More JavaScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.