Default model answers in tool markup; recovery on Opus is over half the spend
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
調査の方向性
Start with internal/session/checkpoint.go around lines 152, 1927, and 3810, then trace the markup retry ladder in internal/session/loop.go around lines 2724-2773. Run the listed unit and org-replication checks on the default model, verify markup recovery stays within the conversation tier or parses the markup, and confirm markreader spend is under 15%.
索引モデルが issue の本文から書いたものです。
説明
Seen on: dev 837b2b0.
Behaviour
On the default model, many turns show the model answered in its own internal markup instead of words · asking again, then gave up after 6 tries …; some managers repeat the ask is not finished · carrying on rather than stopping here. The recovery reader runs on the top model: in one session the markreader role on anthropic/claude-opus-5.5 was $2.29 of $4.27 (54%), and the spend panel read claude-opus-5.5 51%. A fleet on the default cheap model should not spend most of its money on markup recovery.
Replication
- Build dev 837b2b0 (
git checkout 837b2b0 && make build, binarybin/codeaf), or install the dev build withcurl -fsSL https://agentfield.ai/get/devaf | bash. - Use an isolated profile:
export HOME=$(mktemp -d), exportOPENROUTER_API_KEY, and keep the default model (~deepseek/deepseek-v4-flash-latest, crew on auto). - On a busy machine set
task.max_loadto0(/settings, Tasks) so the busy-machine gate does not hold tasks. cd $(mktemp -d) && git init -q && codeaf.- Ask for an org:
Build a small software company for this repo: a product team, a backend team and a qa team, each with a manager, and a global manager over them.Then give each manager one small job. - After 20 to 30 minutes, read
$HOME/.codeaf/v3/usage.jsonland sumspendby role and model (jq -r '[.role,.model,.spend]|@tsv'), and open the spend panel.
Evidence
- Screen:
the model answered in its own internal markup instead of words · asking again,gave up after 6 tries …. - usage.jsonl: role
markreaderonanthropic/claude-opus-5.5= $2.29 of $4.27. internal/session/checkpoint.go:152:roles.Register(roles.RoleMarkReader, roles.TierMastermind)puts the mark reader on the top tier; calls atcheckpoint.go:1927and:3810.internal/session/loop.go:2724-2773: the markup retry ladder (finishing this one on <next>).
Guessed cause
A guess from reading the code, not a confirmed diagnosis. ~deepseek/deepseek-v4-flash-latest on its current lane emits tool markup as text; each such turn wakes the mark reader, which is registered on the mastermind tier, so cheap-model markup failures are paid for at Opus rates.
Acceptance
- e2e: the org replication above on the default model spends under 15% of its total on the
markreaderrole. - Unit: the markup path either parses the model's tool markup or escalates within the conversation's own tier; a structural test pins the mark reader's tier.
Found while writing the public docs; manual text differences are in #1545.
🤖 Generated with Claude Code
- 主要言語
- Go
- スター
- 115
- フォーク
- 14
- 平均マージ
- 9時間 37分
- マージ済み PR(30日)
- 755
環境構築
このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
Agent-Field/CodeAF のほかの issue
-
area:chat bug sev:papercut
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
Agent-Field/CodeAF#1592 ·
メンテナーはふだん 1 日以内に返信
-
area:headless bug sev:critical
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
Agent-Field/CodeAF#1566 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
area:chat feature
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
Agent-Field/CodeAF#1510 ·
メンテナーはふだん 1 日以内に返信
-
area:tests bug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
Agent-Field/CodeAF#1489 ·
メンテナーはふだん 1 日以内に返信
-
area:chat bug good first issue sev:papercut
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
Agent-Field/CodeAF#1470 ·
メンテナーはふだん 1 日以内に返信
Agent-Field/CodeAF の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
rossoctl/context-guru#346 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 90/100
prime-radiant-inc/evener#2883 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
gravitational/teleport#69805 ·
メンテナーはふだん 11 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
Under Poisson sampling, the `PLDAccountant` composes the inner event both before and after samplingオープン
難易度 2/5 半日 初心者へのやさしさ 78/100
google/differential-privacy#496 ·