Default model answers in tool markup; recovery on Opus is over half the spend
Maintainer antworten meist innerhalb von 1 Tag
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 65/100
Rechercherichtung
Start with internal/session/checkpoint.go around lines 152, 1927, and 3810, then trace the markup retry ladder in internal/session/loop.go around lines 2724-2773. Run the listed unit and org-replication checks on the default model, verify markup recovery stays within the conversation tier or parses the markup, and confirm markreader spend is under 15%.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Seen on: dev 837b2b0.
Behaviour
On the default model, many turns show the model answered in its own internal markup instead of words · asking again, then gave up after 6 tries …; some managers repeat the ask is not finished · carrying on rather than stopping here. The recovery reader runs on the top model: in one session the markreader role on anthropic/claude-opus-5.5 was $2.29 of $4.27 (54%), and the spend panel read claude-opus-5.5 51%. A fleet on the default cheap model should not spend most of its money on markup recovery.
Replication
- Build dev 837b2b0 (
git checkout 837b2b0 && make build, binarybin/codeaf), or install the dev build withcurl -fsSL https://agentfield.ai/get/devaf | bash. - Use an isolated profile:
export HOME=$(mktemp -d), exportOPENROUTER_API_KEY, and keep the default model (~deepseek/deepseek-v4-flash-latest, crew on auto). - On a busy machine set
task.max_loadto0(/settings, Tasks) so the busy-machine gate does not hold tasks. cd $(mktemp -d) && git init -q && codeaf.- Ask for an org:
Build a small software company for this repo: a product team, a backend team and a qa team, each with a manager, and a global manager over them.Then give each manager one small job. - After 20 to 30 minutes, read
$HOME/.codeaf/v3/usage.jsonland sumspendby role and model (jq -r '[.role,.model,.spend]|@tsv'), and open the spend panel.
Evidence
- Screen:
the model answered in its own internal markup instead of words · asking again,gave up after 6 tries …. - usage.jsonl: role
markreaderonanthropic/claude-opus-5.5= $2.29 of $4.27. internal/session/checkpoint.go:152:roles.Register(roles.RoleMarkReader, roles.TierMastermind)puts the mark reader on the top tier; calls atcheckpoint.go:1927and:3810.internal/session/loop.go:2724-2773: the markup retry ladder (finishing this one on <next>).
Guessed cause
A guess from reading the code, not a confirmed diagnosis. ~deepseek/deepseek-v4-flash-latest on its current lane emits tool markup as text; each such turn wakes the mark reader, which is registered on the mastermind tier, so cheap-model markup failures are paid for at Opus rates.
Acceptance
- e2e: the org replication above on the default model spends under 15% of its total on the
markreaderrole. - Unit: the markup path either parses the model's tool markup or escalates within the conversation's own tier; a structural test pins the mark reader's tier.
Found while writing the public docs; manual text differences are in #1545.
🤖 Generated with Claude Code
- Vorherrschende Sprache
- Go
- Sterne
- 115
- Forks
- 14
- Ø Merge
- 9 Std. 37 Min.
- Gemergte PRs (30 T.)
- 755
Entwicklungsumgebung
Die Einrichtungsdateien dieses Projekts haben wir noch nicht geprüft. Beginnen Sie mit der README; die allgemeinen Schritte stehen in unserem Leitfaden für den ersten Beitrag.
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus Agent-Field/CodeAF
-
area:chat bug sev:papercut
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 86/100
Agent-Field/CodeAF#1592 ·
Maintainer antworten meist innerhalb von 1 Tag
-
area:headless bug sev:critical
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Agent-Field/CodeAF#1566 · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 1 Tag
-
area:chat feature
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Agent-Field/CodeAF#1510 ·
Maintainer antworten meist innerhalb von 1 Tag
-
area:tests bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
Agent-Field/CodeAF#1489 ·
Maintainer antworten meist innerhalb von 1 Tag
-
area:chat bug good first issue sev:papercut
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
Agent-Field/CodeAF#1470 ·
Maintainer antworten meist innerhalb von 1 Tag
Alle Issues in Agent-Field/CodeAF
Ähnliche Issues
-
Remove CAAPFOffenkind/chore kind/cleanup needs-area
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 86/100
rancher/turtles#2848 · 3 Kommentare ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
Maintainer antworten meist innerhalb von 1 Tag
-
good first issue
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
Maintainer antworten meist innerhalb von 1 Tag
-
priority: low 🌱 type: enhancement 💅🏼
Schwierigkeit 2/5 Ein halber Tag Anfängerfreundlichkeit 84/100
nebari-dev/llm-serving-pack#199 ·
Maintainer antworten meist innerhalb von 3 Tagen
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 90/100
kedacore/keda#8225 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag