Remote/HTTP MCP servers (e.g. atlassian) are stranded 'failed' after every /clear or session relaunch
還沒有人認領這個 Issue。
評估
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 新手友好度
- 45/100
- Issue 類型
- 缺陷
- 描述清晰度
- 基本清楚
- 活躍度
- 活躍
- 技術堆疊
- rust
- 領域
- api, cli, networking
研究方向
使用設定在 ~/.copilot/mcp-config.json 中的 HTTP server 重現故障,然後檢查 MCP reload lock 附近的前景工作階段交接和 copilot_runtime::session::mcp::session_host 記錄。將 /clear 和重新啟動的行為與文件記載的取消序列進行比較。完成標準是:遠端伺服器在任一交接後仍保持連線,或自動恢復連線,而不需要手動重新連線。
由索引模型根據 Issue 內容生成。
描述
Summary
Remote (HTTP) MCP server connections are reliably stranded in a failed state whenever the CLI performs a foreground-session handoff — which happens on every /clear, and apparently on session resume/relaunch too. The MCP connection graph is torn down and rebuilt during the handoff, and slower remote/OAuth-backed servers lose the reconnect race and never recover automatically.
Environment
- CLI version: 1.0.83 (macOS, confirmed already latest via in-app update check)
- MCP server affected:
atlassian(https://mcp.atlassian.com/v2/mcp, HTTP transport with Authorization header) - Other MCP servers configured (
github-mcp-server,codegraph,local-rag— local/stdio) are not visibly affected, likely because their reconnect is fast enough to win the race.
Steps to reproduce
- Start a Copilot CLI session with the
atlassianMCP server configured (~/.copilot/mcp-config.json, HTTP transport). - Confirm it connects fine (
/mcp show atlassian→ connected, tools listed). - Run
/clear(or relaunch/resume a session). - Run
/mcp show atlassianagain.
Expected
atlassian reconnects cleanly like the other MCP servers.
Actual
atlassian is left in:
{
"name": "atlassian",
"status": "failed",
"error": "MCP server \"atlassian\" connection was cancelled"
}
It does not self-heal — it requires a manual reconnect action or a full CLI restart, and even a full restart reproduces the same failure deterministically (confirmed across two separate relaunches, ~20 minutes apart).
Root cause (from ~/.copilot/logs/process-*.log)
Every /clear/relaunch briefly registers a throwaway foreground session, then immediately unregisters it in favor of the real one, within milliseconds:
17:27:22.425Z [INFO] Registering foreground session: 6266dbfb-39da-4797-9265-c37d9655a3fd
17:27:28.717Z [INFO] Unregistering foreground session: 6266dbfb-39da-4797-9265-c37d9655a3fd
17:27:28.720Z [INFO] Registering foreground session: d89d1d9a-99c4-4247-86fe-a75b4da5daad
17:27:28.783Z [INFO] Closing session 6266dbfb-39da-4797-9265-c37d9655a3fd
This handoff triggers a full MCP graph reload while a reload lock is still held from the previous teardown:
17:05:05.167Z [WARNING] [rust:copilot_runtime::session::mcp::session_host] MCP reload lock still held at disposal; draining the graph without it
Every MCP server — including atlassian — gets re-initialized and then almost immediately cancelled mid-handshake:
17:27:26.104Z [INFO] [rust:rmcp::service] Service initialized as client {... "name": "atlassian-mcp-server" ...}
17:27:28.809Z [INFO] [rust:rmcp::service] task cancelled
17:27:28.809Z [INFO] [rust:rmcp::service] serve finished {"quit_reason":"Cancelled"}
Local/stdio servers appear to reconnect fast enough afterward to recover; atlassian's remote HTTP+OAuth handshake is slower and consistently loses the race, leaving it permanently failed with no automatic retry.
Confirmed not a network/auth issue independently:
- DNS resolves fine for
mcp.atlassian.com. curl https://mcp.atlassian.com/v2/mcpreturns401(server reachable, no valid header supplied in the manual test — expected).ATLASSIAN_MCP_AUTH_HEADERenv var is set.
Impact
Every /clear (a very common action) breaks the Atlassian MCP integration and requires a manual reconnect or CLI restart to restore it — significant daily friction for anyone using an HTTP-based MCP server.
Suggested fix
- Don't tear down/reinitialize already-connected MCP servers on a foreground-session handoff that isn't actually changing MCP config.
- If a reload is unavoidable, retry a server that was cancelled mid-handshake instead of leaving it permanently in
failedstate. - Respect the "MCP reload lock still held at disposal" case by waiting for the lock instead of draining the graph without it.
- 主要語言
- Shell
- 星號
- 11.2k
- 分支
- 1.9k
- 平均合併
- 14 小時 16 分鐘
- 30 天內合併 PR
- 6
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
github/copilot-cli 的其他 Issue
-
triage
難度 2/5 1-3 小時 新手友好度 75/100
github/copilot-cli#4932 ·
-
triage
難度 2/5 1-3 小時 新手友好度 78/100
github/copilot-cli#4909 ·
-
triage
難度 2/5 1-3 小時 新手友好度 76/100
github/copilot-cli#4906 ·
-
triage
難度 2/5 1-3 小時 新手友好度 72/100
github/copilot-cli#4848 ·
-
area:agents area:mcp
難度 2/5 1-3 小時 新手友好度 72/100
github/copilot-cli#4729 ·
查看 github/copilot-cli 的全部 Issue
相似的 Issue
-
area: harness bug status: needs-triage
難度 2/5 1-3 小時 新手友好度 75/100
Human-Agent-Society/reef#625 ·
-
module: core
難度 2/5 1-3 小時 新手友好度 75/100
bigbluebutton/bigbluebutton#25849 ·
-
難度 2/5 1-3 小時 新手友好度 75/100
vercel-labs/just-bash#464 ·
-
area:sandbox documentation enhancement platform:linux
難度 1/5 1 小時以內 新手友好度 85/100
anthropics/claude-code#96664 ·
-
cost:cheap severity:medium
難度 2/5 1-3 小時 新手友好度 70/100
fairagro/m4.2_sql_to_arc#227 ·