Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Multi-turn: criteria-derived feedback biases the agent — need a separate, criteria-blind answering prompt (incl. HITL)

Open
#1,228 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
38/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Quiet
Tech stack
typescript
Domain
ai

Research direction

Start by reading packages/shared/src/judge/multi-turn-loop.ts around line 561 and apps/judge/src/feedback-generator.ts, then trace how judgeResults, criteriaRegistry, and HITL tool responses flow through multi-turn runs. Done means the answering path no longer receives judge or criteria context, while also handling coding-agent requests for human input.

Written by the indexing model from the issue text.

Description

author: sinedied

Original author: @sinedied

Summary

In multi-turn runs, the feedback that steers the coding agent between iterations is derived directly from the judge criteria. This means the criteria directly influence the agent's output — the run is effectively "taught to the test," which biases/contaminates the evaluation.

Why it matters

Criteria are supposed to measure the agent, not steer it. When criteria-derived feedback is fed back as the next instruction, a criterion both defines success and tells the agent how to achieve it, inflating scores and invalidating cross-config comparisons (e.g. baseline vs. with-skills).

Current behavior (code)

  • packages/shared/src/judge/multi-turn-loop.ts:561 — nextPrompt = judgeFeedback; The coding agent's next-turn prompt is the judge feedback.
  • apps/judge/src/feedback-generator.ts — feedback is generated from judgeResults + criteriaRegistry. It masks the wording (instructed not to mention "evaluation/judge/criteria") but the content is still criteria-derived, so the steering/bias remains.
  • There is no separate prompt governing how the agent should proceed in multi-turn independent of the criteria.

Proposal

  • Introduce a separate multi-turn "answering"/continuation prompt that instructs the agent how to proceed across turns, without any access to the judge/criteria context.
  • Keep the judge/criteria strictly on the measurement side; the answering agent must be criteria-blind to avoid biasing outputs.

HITL gap

  • The same criteria-blind answering path should handle responding to the coding agent's tool response for human-in-the-loop (HITL) calls (e.g. when the agent asks a clarifying question / requests input mid-task). This doesn't appear to be handled today — the loop only feeds criteria-derived feedback as the next prompt.

Related (not duplicates)

  • #162 — feedback LLM determinism (different concern).
  • #679 — multi-turn success reporting (different concern).
Dominant language
TypeScript
Stars
5
Forks
5
Avg merge
3d 5h
Merged PRs (30d)
26

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/scope

All issues in microsoft/scope

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.