Default model answers in tool markup; recovery on Opus is over half the spend
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 65/100
Direzione di ricerca
Start with internal/session/checkpoint.go around lines 152, 1927, and 3810, then trace the markup retry ladder in internal/session/loop.go around lines 2724-2773. Run the listed unit and org-replication checks on the default model, verify markup recovery stays within the conversation tier or parses the markup, and confirm markreader spend is under 15%.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Seen on: dev 837b2b0.
Behaviour
On the default model, many turns show the model answered in its own internal markup instead of words · asking again, then gave up after 6 tries …; some managers repeat the ask is not finished · carrying on rather than stopping here. The recovery reader runs on the top model: in one session the markreader role on anthropic/claude-opus-5.5 was $2.29 of $4.27 (54%), and the spend panel read claude-opus-5.5 51%. A fleet on the default cheap model should not spend most of its money on markup recovery.
Replication
- Build dev 837b2b0 (
git checkout 837b2b0 && make build, binarybin/codeaf), or install the dev build withcurl -fsSL https://agentfield.ai/get/devaf | bash. - Use an isolated profile:
export HOME=$(mktemp -d), exportOPENROUTER_API_KEY, and keep the default model (~deepseek/deepseek-v4-flash-latest, crew on auto). - On a busy machine set
task.max_loadto0(/settings, Tasks) so the busy-machine gate does not hold tasks. cd $(mktemp -d) && git init -q && codeaf.- Ask for an org:
Build a small software company for this repo: a product team, a backend team and a qa team, each with a manager, and a global manager over them.Then give each manager one small job. - After 20 to 30 minutes, read
$HOME/.codeaf/v3/usage.jsonland sumspendby role and model (jq -r '[.role,.model,.spend]|@tsv'), and open the spend panel.
Evidence
- Screen:
the model answered in its own internal markup instead of words · asking again,gave up after 6 tries …. - usage.jsonl: role
markreaderonanthropic/claude-opus-5.5= $2.29 of $4.27. internal/session/checkpoint.go:152:roles.Register(roles.RoleMarkReader, roles.TierMastermind)puts the mark reader on the top tier; calls atcheckpoint.go:1927and:3810.internal/session/loop.go:2724-2773: the markup retry ladder (finishing this one on <next>).
Guessed cause
A guess from reading the code, not a confirmed diagnosis. ~deepseek/deepseek-v4-flash-latest on its current lane emits tool markup as text; each such turn wakes the mark reader, which is registered on the mastermind tier, so cheap-model markup failures are paid for at Opus rates.
Acceptance
- e2e: the org replication above on the default model spends under 15% of its total on the
markreaderrole. - Unit: the markup path either parses the model's tool markup or escalates within the conversation's own tier; a structural test pins the mark reader's tier.
Found while writing the public docs; manual text differences are in #1545.
🤖 Generated with Claude Code
- Lingua principale
- Go
- Stelle
- 115
- Fork
- 14
- Merge medio
- 9h 37m
- PR unite (30g)
- 755
Preparare l'ambiente
Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Agent-Field/CodeAF
-
area:chat bug sev:papercut
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
Agent-Field/CodeAF#1592 ·
I maintainer di solito rispondono entro 1 giorno
-
area:headless bug sev:critical
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
Agent-Field/CodeAF#1566 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
area:chat feature
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
Agent-Field/CodeAF#1510 ·
I maintainer di solito rispondono entro 1 giorno
-
area:tests bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
Agent-Field/CodeAF#1489 ·
I maintainer di solito rispondono entro 1 giorno
-
area:chat bug good first issue sev:papercut
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
Agent-Field/CodeAF#1470 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di Agent-Field/CodeAF
Issue simili
-
agentic-workflows
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 1 giorno
-
priority/4/normal status/needs-triage type/bug/unconfirmed
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
authelia/authelia#13292 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
blinklabs-io/actions#138 ·
I maintainer di solito rispondono entro 1 giorno
-
[UI] AlbumDetails collapses multi-genre list to single primary genre on viewports < lg breakpointAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno