Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Default model answers in tool markup; recovery on Opus is over half the spend

クローズ
#1,570 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
65/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
go
領域
ai, backend, cli

調査の方向性

Start with internal/session/checkpoint.go around lines 152, 1927, and 3810, then trace the markup retry ladder in internal/session/loop.go around lines 2724-2773. Run the listed unit and org-replication checks on the default model, verify markup recovery stays within the conversation tier or parses the markup, and confirm markreader spend is under 15%.

索引モデルが issue の本文から書いたものです。

説明

area:provider bug sev:critical

Seen on: dev 837b2b0.

Behaviour

On the default model, many turns show the model answered in its own internal markup instead of words · asking again, then gave up after 6 tries …; some managers repeat the ask is not finished · carrying on rather than stopping here. The recovery reader runs on the top model: in one session the markreader role on anthropic/claude-opus-5.5 was $2.29 of $4.27 (54%), and the spend panel read claude-opus-5.5 51%. A fleet on the default cheap model should not spend most of its money on markup recovery.

Replication

  1. Build dev 837b2b0 (git checkout 837b2b0 && make build, binary bin/codeaf), or install the dev build with curl -fsSL https://agentfield.ai/get/devaf | bash.
  2. Use an isolated profile: export HOME=$(mktemp -d), export OPENROUTER_API_KEY, and keep the default model (~deepseek/deepseek-v4-flash-latest, crew on auto).
  3. On a busy machine set task.max_load to 0 (/settings, Tasks) so the busy-machine gate does not hold tasks.
  4. cd $(mktemp -d) && git init -q && codeaf.
  5. Ask for an org: Build a small software company for this repo: a product team, a backend team and a qa team, each with a manager, and a global manager over them. Then give each manager one small job.
  6. After 20 to 30 minutes, read $HOME/.codeaf/v3/usage.jsonl and sum spend by role and model (jq -r '[.role,.model,.spend]|@tsv'), and open the spend panel.

Evidence

  • Screen: the model answered in its own internal markup instead of words · asking again, gave up after 6 tries ….
  • usage.jsonl: role markreader on anthropic/claude-opus-5.5 = $2.29 of $4.27.
  • internal/session/checkpoint.go:152: roles.Register(roles.RoleMarkReader, roles.TierMastermind) puts the mark reader on the top tier; calls at checkpoint.go:1927 and :3810.
  • internal/session/loop.go:2724-2773: the markup retry ladder (finishing this one on <next>).

Guessed cause

A guess from reading the code, not a confirmed diagnosis. ~deepseek/deepseek-v4-flash-latest on its current lane emits tool markup as text; each such turn wakes the mark reader, which is registered on the mastermind tier, so cheap-model markup failures are paid for at Opus rates.

Acceptance

  • e2e: the org replication above on the default model spends under 15% of its total on the markreader role.
  • Unit: the markup path either parses the model's tool markup or escalates within the conversation's own tier; a structural test pins the mark reader's tier.

Found while writing the public docs; manual text differences are in #1545.

🤖 Generated with Claude Code

主要言語
Go
スター
115
フォーク
14
平均マージ
9時間 37分
マージ済み PR(30日)
755

環境構築

このプロジェクトの環境構築ファイルはまだ確認していません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

Agent-Field/CodeAF のほかの issue

Agent-Field/CodeAF の issue をすべて見る

似ている issue

Go の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。