Multi-turn: criteria-derived feedback biases the agent — need a separate, criteria-blind answering prompt (incl. HITL)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 38/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Tranquilla
- Stack tecnologico
- typescript
- Ambito
- ai
Direzione di ricerca
Start by reading packages/shared/src/judge/multi-turn-loop.ts around line 561 and apps/judge/src/feedback-generator.ts, then trace how judgeResults, criteriaRegistry, and HITL tool responses flow through multi-turn runs. Done means the answering path no longer receives judge or criteria context, while also handling coding-agent requests for human input.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Original author: @sinedied
Summary
In multi-turn runs, the feedback that steers the coding agent between iterations is derived directly from the judge criteria. This means the criteria directly influence the agent's output — the run is effectively "taught to the test," which biases/contaminates the evaluation.
Why it matters
Criteria are supposed to measure the agent, not steer it. When criteria-derived feedback is fed back as the next instruction, a criterion both defines success and tells the agent how to achieve it, inflating scores and invalidating cross-config comparisons (e.g. baseline vs. with-skills).
Current behavior (code)
packages/shared/src/judge/multi-turn-loop.ts:561—nextPrompt = judgeFeedback;The coding agent's next-turn prompt is the judge feedback.apps/judge/src/feedback-generator.ts— feedback is generated fromjudgeResults+criteriaRegistry. It masks the wording (instructed not to mention "evaluation/judge/criteria") but the content is still criteria-derived, so the steering/bias remains.- There is no separate prompt governing how the agent should proceed in multi-turn independent of the criteria.
Proposal
- Introduce a separate multi-turn "answering"/continuation prompt that instructs the agent how to proceed across turns, without any access to the judge/criteria context.
- Keep the judge/criteria strictly on the measurement side; the answering agent must be criteria-blind to avoid biasing outputs.
HITL gap
- The same criteria-blind answering path should handle responding to the coding agent's tool response for human-in-the-loop (HITL) calls (e.g. when the agent asks a clarifying question / requests input mid-task). This doesn't appear to be handled today — the loop only feeds criteria-derived feedback as the next prompt.
Related (not duplicates)
- #162 — feedback LLM determinism (different concern).
- #679 — multi-turn success reporting (different concern).
- Lingua principale
- TypeScript
- Stelle
- 7
- Fork
- 13
- Merge medio
- 3g 5h
- PR unite (30g)
- 26
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/scope
-
type: worker-update
Difficoltà 1/5 1-3 ore Idoneità per principianti 85/100
I maintainer di solito rispondono entro 1 giorno
-
type: worker-update
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
type: worker-update
Difficoltà 1/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
Clarify that prompt features only categorize prompts, don't impact runsForse già presa @DerrickUnleashed l’ha presa 2 giorni fa. Apertaauthor: JaGord documentation good first issue UI
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Submit Run task prompt picker should only search requirement promptsForse già presa @cedricvidal l’ha presa 68 giorni fa. Apertaauthor: cedricvidal bug portal
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di microsoft/scope
Issue simili
-
dx hacktoberfest help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
I maintainer di solito rispondono entro 1 giorno
-
documentation
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
cloudflare/agents#2498 ·
I maintainer di solito rispondono entro 1 giorno
-
Missing repro Platform: Android
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
software-mansion/react-native-reanimated#10816 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
e2e-failure ready-to-code
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
redhat-developer/rhdh-plugin-export-overlays#4129 ·
I maintainer di solito rispondono entro 1 giorno