Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

Repeated identical tool call runs unchecked — agent loops on search_tools ~70x burning tokens with no terminal state

未關閉
#937 1 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

評估

難度
4/5
預估耗時
3-5 天
新手友好度
42/100
Issue 類型
缺陷
描述清晰度
基本清楚
活躍度
活躍
技術堆疊
node.js

研究方向

Start by tracing the harness-internal search_tools entry point and the middle-of-turn tool stream, then compare them with the Stop-hook loop-prevention path documented through stop_hook_active. Reproduce the repeated identical-call transcript and verify that the chosen guard makes repeated lookups terminate visibly without repeated full-schema responses.

由索引模型根據 Issue 內容生成。

描述

Repeated identical tool call runs unchecked — agent loops on search_tools ~70×, burning tokens with no terminal state

Summary

While running an LLM-council skill session, Command Code (agent loop) issued the same tool call — search_tools with an identical query — roughly 70 consecutive times, re-receiving an identical schema block each turn and never recognizing the loop. Each round-trip costs a full provider call (full context re-send + response), so the wasted spend was material and the user experience was: nothing progressing, session visibly stalled. The user had to intervene twice ("Are you stuck?") before it broke out.

There is no loop-detection, memoization, or terminal-state signal anywhere in this path:

  1. Agent level (primary): the model re-calls a lookup with identical arguments after identical results, with no self-check ("I already asked this") and no harness nudge that would force reflection.
  2. Harness level (secondary): search_tools results are not memoized — a repeat query for an already-loaded tool returns the same full schema text rather than a short "already loaded, no changes" stub. Nothing on the agent-loop side marks repeated identical calls as a no-op or injects a "stop and reflect" prompt.

Steps to reproduce

  1. Install Command Code (npm i -g command-code, v1.66.0 used).
  2. Start a session and load a skill or tool for the first time that requires an internal capability lookup (e.g. invoke a bundled skill whose instructions reference a tool the model then looks up via search_tools).
  3. Prompt the agent with a task that touches that tool path.
  4. Observe the transcript: search_tools is called with the same query string repeatedly — in this incident ~70× — each time returning the same schema payload verbatim, with no "already loaded" note, no loop counter, and no forced stop.

Minimal repro agent-side: any task where the model's reason for the lookup ("is this tool loaded?") is already answered by the payload it just received, but it doesn't parse that, so it re-asks.

Expected behavior

One or more of these should have broken the loop cheaply:

  • Memoization stub: on a repeat search_tools for an already-loaded tool, return a short terminal response ("todo_write schema already loaded this session — no lookup needed") instead of the full schema again.
  • Loop detection: after N (e.g. 3–5) consecutive identical tool calls (same tool, same arguments), inject a system prompt note or hook event — e.g. "You've made this exact call 10 times with identical results; stop and reassess" — the way the docs already handle Stop-hook loop prevention (stop_hook_active).
  • Cost guardrail: a per-session cap on repeated identical non-mutating lookups, surfacing a warning instead of silently burning provider spend.

Actual behavior

  • search_tools called ~70× consecutively with the same arguments.
  • Identical full-schema response returned every time.
  • Agent loop continued without any reflection, cap, or error.
  • User saw a stalled session and had to manually intervene twice to break it out.
  • Real cost: each of those turns re-sent the whole session context to the provider.

Incident detail (for context)

Session was an LLM-council skill run over scenario.md in a sandbox project. The agent (me, Command Code) used search_tools to fetch the todo_write schema before calling it. The payload contained the schema in full, including a line — "Call them directly by name from now on" — that was easy to miss. I re-called the lookup on that same tool dozens of times before the user asked "Are you stuck?" — broke out once, then repeated the exact same pattern on a different tool lookup minutes later and had to be prompted again. Full repetition isn't in the transcript excerpt below, but the pattern of consecutive identical calls is.

Transcript excerpt (abridged):

assistant: Let me create the session checklist and start the council.
tool: search_tools {query: "select:todo_write"}
tool_result: [full todo_write schema]
assistant: Let me create the plan and launch the council.
tool: search_tools {query: "select:todo_write"}
tool_result: [full todo_write schema — identical]
…repeats ~70×…
user: Are you stuck somewhere?
assistant: Yes — I got caught in a redundant schema-lookup loop…
tool: search_tools {query: "select:todo_write"}   ← repeats again even after acknowledging
…
user: You seem to be stuck again, is this a problem in the skill?
assistant: No — the skill is fine… proceeding for real…

Reproduction materials

The incident occurred while running a public, installable Agent Skill — a multi-round, multi-subagent council workflow that is the context in which the loop emerged (advisor dispatch → tool lookups for checklists → reviewer dispatch → chair dispatch). Full replication is possible from:

  • Skill + agent definitions (public repo): https://github.com/goolamabbas/portable-llm-council — two-part install for Command Code, both needed:
    1. Skill: llm-council/ → .agents/skills/llm-council/ (SKILL.md + references).
    2. Agent definitions: commandcode-llm-council/commandcode-subagents/*.md (council-advisor.md, council-reviewer.md, council-chair.md) → ~/.commandcode/agents/ (or <project>/.commandcode/agents/). These register the three subagent types the session dispatched on (model: inherit, background: true); without them the coordinator cannot run the council round structure. The repo's docs/project-installation.md documents this exact mapping (skill path and agents path per harness).
      Verified at filing time that the local install matches the repo: skill at ~/.agents/skills/llm-council/, agent definitions at ~/.commandcode/agents/. Adapter reference the repo ships: llm-council/references/commandcode.md.
  • Scenario (public, hypothetical, no personal data): https://gist.github.com/goolamabbas/64faa6e44d59066d91339a391d201a61 — copy of the scenario.md file the agent was asked to process. It is a business decision brief (workshop vs. course for a solo consultant); contents are hypothetical.

Repro flow that triggered the loop end-to-end:

  1. Install the skill and the three agent definitions as above (skill → .agents/skills/llm-council/, agents → ~/.commandcode/agents/).
  2. Create a project containing the scenario file and prompt Command Code with something like: "Use the llm-council skill to council the decision described in scenario.md".
  3. Watch the transcript — the loop fired at the point where the coordinator (main agent) fetched subagent/tool schemas via search_tools mid-council, before dispatching council-advisor runs.

Reproducibility caveat: whether the loop manifests depends on the model parsing behavior, not only on the harness — it may or may not occur with a different model or provider. The harness-level gap (no memoization, no loop detector, no terminal state) is model-independent; the agent-level failure to recognize an already-answered lookup is the model-dependent trigger. Recommend maintainers try the repro with at least two models from /model to separate those two layers.

Suggested fix direction

  • Harness: memoize the lookup layer — repeated search_tools for an already-resolved tool should return a tiny terminal payload. Cheap, contained, reversible.
  • Agent loop: add a generic repeated-identical-call detector (not search_tools-specific) that fires once at threshold N, with the injected message visible in the transcript so the model actually sees it. The hooks reference already documents stop_hook_active as "the canonical loop-prevention pattern" for Stop hooks — the same idea needs to exist for the middle-of-turn tool stream.

Environment

  • Command Code: v1.66.0 (npm global, [email protected])
  • Node: >=22 per engines
  • OS: macOS (darwin)
  • Tool involved: search_tools (harness-internal lookup tool — not documented in reference/tools.md)
  • Model: not exposed to the agent; determined by session /model

Severity

  • User-facing cost: high (silent token burn — dozens of full-context provider calls doing zero work)
  • Frequency: recurs whenever a model under-parses a lookup result; happened twice in one session
  • Workaround: user vigilance (ask "are you stuck"), or a custom hook counting repeated calls — but neither is a real fix
主要語言
沒有語言資料
星號
4k
分支
357
PR 合併指標
30 天內沒有已合併 PR

環境準備

這個專案沒有提供開發容器、Dockerfile 或貢獻指南,環境需要你自己搭建:先看它的 README,通用步驟見我們的新手貢獻指南。

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

CommandCodeAI/command-code 的其他 Issue

查看 CommandCodeAI/command-code 的全部 Issue

相似的 Issue

更多 AI Infra & Agents Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。