Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Add benchmark suite comparing native policy learning vs fine-tuning on representative agent workloads

Open
#1 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 4 days

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
azure, python

Research direction

Start from the existing Python learning loop: its softmax policy over discrete actions, Azure AI Evaluation evaluators, REINFORCE updates, and persisted episodes. Build the benchmark around at least two representative task families, compare against a weight fine-tuning baseline using shared reward and success criteria, and make one command reproduce seeded results plus the bounded-claim documentation.

Written by the indexing model from the issue text.

Description

The Toulmin analysis calls out that the strongest claim in the whitepaper is the architectural case for native policy learning, but the weakest point is the lack of benchmark evidence showing that policy-layer learning achieves comparable quality gains on real workloads. This is the highest-value gap to close because it is the direct test of the core substitution claim.

The current repo already models the learning loop as a softmax policy over discrete actions, judged by three Azure AI Evaluation evaluators, with REINFORCE updates and persisted episodes. The missing piece is a hard benchmark that compares this approach against a weight fine-tuning baseline on at least one realistic task family, using the same reward and success criteria.

Suggested scope

Add benchmark harnesses for at least two or three representative task families, such as structured tool use, RAG, and prompt/routing optimization.
Compare native policy learning against a fine-tuning baseline on both reward and task-success metrics.
Publish a reproducible results table with seeds, evaluation protocol, and acceptance criteria.
Include a clear “bounded claim” section in docs that explains when policy learning is expected to work and when it is not.

Acceptance criteria

Benchmark results are reproducible from a single command or script.
Results include both reward and task-success outcomes.
The comparison explicitly calls out the conditions under which native policy learning is expected to match or fall short of fine-tuning.

Dominant language
Python
Stars
10
Forks
9
Avg merge
2d 5h
Merged PRs (30d)
10

Getting set up

We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/agent-learning

All issues in microsoft/agent-learning

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.