Repeated identical tool call runs unchecked — agent loops on search_tools ~70x burning tokens with no terminal state
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 42/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- node.js
- Lĩnh vực
- ai-infra-agents, cli
Hướng nghiên cứu
Start by tracing the harness-internal search_tools entry point and the middle-of-turn tool stream, then compare them with the Stop-hook loop-prevention path documented through stop_hook_active. Reproduce the repeated identical-call transcript and verify that the chosen guard makes repeated lookups terminate visibly without repeated full-schema responses.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Repeated identical tool call runs unchecked — agent loops on search_tools ~70×, burning tokens with no terminal state
Summary
While running an LLM-council skill session, Command Code (agent loop) issued the same tool call — search_tools with an identical query — roughly 70 consecutive times, re-receiving an identical schema block each turn and never recognizing the loop. Each round-trip costs a full provider call (full context re-send + response), so the wasted spend was material and the user experience was: nothing progressing, session visibly stalled. The user had to intervene twice ("Are you stuck?") before it broke out.
There is no loop-detection, memoization, or terminal-state signal anywhere in this path:
- Agent level (primary): the model re-calls a lookup with identical arguments after identical results, with no self-check ("I already asked this") and no harness nudge that would force reflection.
- Harness level (secondary):
search_toolsresults are not memoized — a repeat query for an already-loaded tool returns the same full schema text rather than a short "already loaded, no changes" stub. Nothing on the agent-loop side marks repeated identical calls as a no-op or injects a "stop and reflect" prompt.
Steps to reproduce
- Install Command Code (
npm i -g command-code, v1.66.0 used). - Start a session and load a skill or tool for the first time that requires an internal capability lookup (e.g. invoke a bundled skill whose instructions reference a tool the model then looks up via
search_tools). - Prompt the agent with a task that touches that tool path.
- Observe the transcript:
search_toolsis called with the same query string repeatedly — in this incident ~70× — each time returning the same schema payload verbatim, with no "already loaded" note, no loop counter, and no forced stop.
Minimal repro agent-side: any task where the model's reason for the lookup ("is this tool loaded?") is already answered by the payload it just received, but it doesn't parse that, so it re-asks.
Expected behavior
One or more of these should have broken the loop cheaply:
- Memoization stub: on a repeat
search_toolsfor an already-loaded tool, return a short terminal response ("todo_writeschema already loaded this session — no lookup needed") instead of the full schema again. - Loop detection: after N (e.g. 3–5) consecutive identical tool calls (same tool, same arguments), inject a system prompt note or hook event — e.g. "You've made this exact call 10 times with identical results; stop and reassess" — the way the docs already handle Stop-hook loop prevention (
stop_hook_active). - Cost guardrail: a per-session cap on repeated identical non-mutating lookups, surfacing a warning instead of silently burning provider spend.
Actual behavior
search_toolscalled ~70× consecutively with the same arguments.- Identical full-schema response returned every time.
- Agent loop continued without any reflection, cap, or error.
- User saw a stalled session and had to manually intervene twice to break it out.
- Real cost: each of those turns re-sent the whole session context to the provider.
Incident detail (for context)
Session was an LLM-council skill run over scenario.md in a sandbox project. The agent (me, Command Code) used search_tools to fetch the todo_write schema before calling it. The payload contained the schema in full, including a line — "Call them directly by name from now on" — that was easy to miss. I re-called the lookup on that same tool dozens of times before the user asked "Are you stuck?" — broke out once, then repeated the exact same pattern on a different tool lookup minutes later and had to be prompted again. Full repetition isn't in the transcript excerpt below, but the pattern of consecutive identical calls is.
Transcript excerpt (abridged):
assistant: Let me create the session checklist and start the council.
tool: search_tools {query: "select:todo_write"}
tool_result: [full todo_write schema]
assistant: Let me create the plan and launch the council.
tool: search_tools {query: "select:todo_write"}
tool_result: [full todo_write schema — identical]
…repeats ~70×…
user: Are you stuck somewhere?
assistant: Yes — I got caught in a redundant schema-lookup loop…
tool: search_tools {query: "select:todo_write"} ← repeats again even after acknowledging
…
user: You seem to be stuck again, is this a problem in the skill?
assistant: No — the skill is fine… proceeding for real…
Reproduction materials
The incident occurred while running a public, installable Agent Skill — a multi-round, multi-subagent council workflow that is the context in which the loop emerged (advisor dispatch → tool lookups for checklists → reviewer dispatch → chair dispatch). Full replication is possible from:
- Skill + agent definitions (public repo): https://github.com/goolamabbas/portable-llm-council — two-part install for Command Code, both needed:
- Skill:
llm-council/→.agents/skills/llm-council/(SKILL.md + references). - Agent definitions:
commandcode-llm-council/commandcode-subagents/*.md(council-advisor.md,council-reviewer.md,council-chair.md) →~/.commandcode/agents/(or<project>/.commandcode/agents/). These register the three subagent types the session dispatched on (model: inherit,background: true); without them the coordinator cannot run the council round structure. The repo'sdocs/project-installation.mddocuments this exact mapping (skill path and agents path per harness).
Verified at filing time that the local install matches the repo: skill at~/.agents/skills/llm-council/, agent definitions at~/.commandcode/agents/. Adapter reference the repo ships:llm-council/references/commandcode.md.
- Skill:
- Scenario (public, hypothetical, no personal data): https://gist.github.com/goolamabbas/64faa6e44d59066d91339a391d201a61 — copy of the
scenario.mdfile the agent was asked to process. It is a business decision brief (workshop vs. course for a solo consultant); contents are hypothetical.
Repro flow that triggered the loop end-to-end:
- Install the skill and the three agent definitions as above (skill →
.agents/skills/llm-council/, agents →~/.commandcode/agents/). - Create a project containing the scenario file and prompt Command Code with something like: "Use the llm-council skill to council the decision described in scenario.md".
- Watch the transcript — the loop fired at the point where the coordinator (main agent) fetched subagent/tool schemas via
search_toolsmid-council, before dispatchingcouncil-advisorruns.
Reproducibility caveat: whether the loop manifests depends on the model parsing behavior, not only on the harness — it may or may not occur with a different model or provider. The harness-level gap (no memoization, no loop detector, no terminal state) is model-independent; the agent-level failure to recognize an already-answered lookup is the model-dependent trigger. Recommend maintainers try the repro with at least two models from /model to separate those two layers.
Suggested fix direction
- Harness: memoize the lookup layer — repeated
search_toolsfor an already-resolved tool should return a tiny terminal payload. Cheap, contained, reversible. - Agent loop: add a generic repeated-identical-call detector (not
search_tools-specific) that fires once at threshold N, with the injected message visible in the transcript so the model actually sees it. The hooks reference already documentsstop_hook_activeas "the canonical loop-prevention pattern" for Stop hooks — the same idea needs to exist for the middle-of-turn tool stream.
Environment
- Command Code: v1.66.0 (npm global,
[email protected]) - Node: >=22 per engines
- OS: macOS (darwin)
- Tool involved:
search_tools(harness-internal lookup tool — not documented inreference/tools.md) - Model: not exposed to the agent; determined by session
/model
Severity
- User-facing cost: high (silent token burn — dozens of full-context provider calls doing zero work)
- Frequency: recurs whenever a model under-parses a lookup result; happened twice in one session
- Workaround: user vigilance (ask "are you stuck"), or a custom hook counting repeated calls — but neither is a real fix
- Ngôn ngữ chính
- Không có dữ liệu ngôn ngữ
- Star
- 4k
- Fork
- 357
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của CommandCodeAI/command-code
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
CommandCodeAI/command-code#933 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
CommandCodeAI/command-code#855 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
CommandCodeAI/command-code#841 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
CommandCodeAI/command-code#655 · 1 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
CommandCodeAI/command-code#608 ·
Tất cả issue của CommandCodeAI/command-code
Issue tương tự
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
mrutunjay-kinagi/ragsearch#129 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
add latest sol modelĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
netlify-labs/nax#56 · 2 bình luận · 1 reaction ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
msgbyte/dao-browser#64 ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
langchain-ai/langgraphjs#2916 ·
Maintainer thường phản hồi trong vòng 1 ngày