System prompt labels shell commands as tools, causing invalid tool calls and autopilot cost
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 68/100
Direzione di ricerca
Cerca nel repository il testo esatto Available tools: git, curl, gh e traccia il punto di ingresso del prompt di sistema di init. Verifica che il prompt risultante distingua i comandi shell dagli strumenti richiamabili, quindi esamina gli eventuali test esistenti del prompt o della CLI per verificarne la copertura.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Describe the bug
In Copilot CLI non-interactive/autopilot sessions, the init system prompt labels shell commands as "Available tools":
<environment_context>
...
* Available tools: git, curl, gh
</environment_context>
In the observed session, curl was available as a shell command, but not as a callable Copilot tool/function. The model treated it as a callable tool and emitted:
tool= curl args= {'url': 'https://app.notion.com/p/...'}
The runtime then returned:
Tool 'curl' does not exist.
This looks like prompt-surface ambiguity rather than only a model mistake: the system prompt uses the word "tools" for shell commands while callable tools are also exposed to the model as tools.
Related schema confusion observed in the same testing work
The same ambiguity showed up around other tool workflows:
- Shell/session tools were called without the required prior shell state:
Multiple validation errors:
- "shellId": Required
- "delay": Required
- Sub-agent/task invocation was attempted with an incomplete schema:
{name: "task", arguments: {agent_type: "translation-validator", description: "Validate changelog files", prompt: "...", mode: "background"}}
Runtime response:
"name": Required
The model then switched to background agent + read_agent, which worked but added extra turns and runtime cost.
Why this matters
Autopilot mode absorbs these validation errors and keeps going, but each invalid tool call costs extra model requests, tokens, and wall time. This is especially visible with smaller/local models that are more sensitive to tool-surface ambiguity.
In our prompt optimization test, making the prompt explicitly avoid the ambiguous paths removed the validation errors across 10/10 successful runs:
- Require direct file write, not shell input/bash write
- Require
translation-validatorwithname=translation-validator - Require synchronous validator, not background/read_agent
- Require validator to read only output files, not Notion/web/curl
- Use
events.jsonlsession.task_completeas completion signal, not stdout
Affected version
Observed in Copilot CLI session metadata:
copilotVersion: 1.0.39
model: Qwen3.6-35B-A3B-bf16
mode: --autopilot, non-interactive
Expected behavior
The init system prompt should clearly distinguish callable tools from shell commands. For example:
* Available shell commands: git, curl, gh
instead of:
* Available tools: git, curl, gh
It would also help if the tool guidance made preconditions and required fields harder to miss, especially for:
- shell/session output tools that require
shellId task/ sub-agent tools that requirename- background agent flows that require
read_agent
Suggested fixes
- Rename
Available tools: git, curl, ghtoAvailable shell commands: git, curl, ghin the system prompt. - Explicitly state that shell commands must be run through the bash/shell tool and are not callable tool names.
- Add stronger schema guidance for
task/ sub-agent invocation, including requirednameand when not to use background/read_agent. - Add stronger precondition wording for shell output/input tools that require an existing
shellId.
Additional context
This is related in impact to task completion/output reliability issues, but it is a distinct problem: tool-surface ambiguity in the init prompt causes invalid tool calls and unnecessary autopilot cost.
- Lingua principale
- Shell
- Stelle
- 11.2k
- Fork
- 1.9k
- Merge medio
- 14h 16m
- PR unite (30g)
- 6
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di github/copilot-cli
-
triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
github/copilot-cli#4848 ·
-
area:agents area:mcp
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
github/copilot-cli#4729 ·
-
area:sessions
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
github/copilot-cli#4712 ·
-
triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
github/copilot-cli#4638 ·
-
Expose large_output_file_path on TaskShellProgress so clients can read complete shell-task output Apertaarea:tools
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
github/copilot-cli#4630 · 1 commento ·
Tutte le issue di github/copilot-cli
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
community-scripts/ProxmoxVE#17425 · 1 commento ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 90/100
danielmiessler/LifeOS#2218 ·
-
docs(agents): strengthen the no-backslash-escaped-backticks rule with an issue-creation example Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
-
technical-debt
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
ll7/robot_sf_ll7#9560 ·
-
Update ghgrab to 2.1.0 Apertapackage-update
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
oSoWoSo/vOid_Community_repOsitory#148 · 1 commento ·