Remote/HTTP MCP servers (e.g. atlassian) are stranded 'failed' after every /clear or session relaunch
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 45/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- rust
- Lĩnh vực
- api, cli, networking
Hướng nghiên cứu
Tái hiện lỗi với một HTTP server được cấu hình trong ~/.copilot/mcp-config.json, sau đó kiểm tra các log về việc chuyển giao foreground session và copilot_runtime::session::mcp::session_host xung quanh MCP reload lock. So sánh hành vi của /clear và việc khởi chạy lại với chuỗi hủy đã được ghi lại trong tài liệu. Hoàn tất khi remote server vẫn được kết nối hoặc tự động khôi phục sau một trong hai lần chuyển giao mà không cần reconnect thủ công.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
Remote (HTTP) MCP server connections are reliably stranded in a failed state whenever the CLI performs a foreground-session handoff — which happens on every /clear, and apparently on session resume/relaunch too. The MCP connection graph is torn down and rebuilt during the handoff, and slower remote/OAuth-backed servers lose the reconnect race and never recover automatically.
Environment
- CLI version: 1.0.83 (macOS, confirmed already latest via in-app update check)
- MCP server affected:
atlassian(https://mcp.atlassian.com/v2/mcp, HTTP transport with Authorization header) - Other MCP servers configured (
github-mcp-server,codegraph,local-rag— local/stdio) are not visibly affected, likely because their reconnect is fast enough to win the race.
Steps to reproduce
- Start a Copilot CLI session with the
atlassianMCP server configured (~/.copilot/mcp-config.json, HTTP transport). - Confirm it connects fine (
/mcp show atlassian→ connected, tools listed). - Run
/clear(or relaunch/resume a session). - Run
/mcp show atlassianagain.
Expected
atlassian reconnects cleanly like the other MCP servers.
Actual
atlassian is left in:
{
"name": "atlassian",
"status": "failed",
"error": "MCP server \"atlassian\" connection was cancelled"
}
It does not self-heal — it requires a manual reconnect action or a full CLI restart, and even a full restart reproduces the same failure deterministically (confirmed across two separate relaunches, ~20 minutes apart).
Root cause (from ~/.copilot/logs/process-*.log)
Every /clear/relaunch briefly registers a throwaway foreground session, then immediately unregisters it in favor of the real one, within milliseconds:
17:27:22.425Z [INFO] Registering foreground session: 6266dbfb-39da-4797-9265-c37d9655a3fd
17:27:28.717Z [INFO] Unregistering foreground session: 6266dbfb-39da-4797-9265-c37d9655a3fd
17:27:28.720Z [INFO] Registering foreground session: d89d1d9a-99c4-4247-86fe-a75b4da5daad
17:27:28.783Z [INFO] Closing session 6266dbfb-39da-4797-9265-c37d9655a3fd
This handoff triggers a full MCP graph reload while a reload lock is still held from the previous teardown:
17:05:05.167Z [WARNING] [rust:copilot_runtime::session::mcp::session_host] MCP reload lock still held at disposal; draining the graph without it
Every MCP server — including atlassian — gets re-initialized and then almost immediately cancelled mid-handshake:
17:27:26.104Z [INFO] [rust:rmcp::service] Service initialized as client {... "name": "atlassian-mcp-server" ...}
17:27:28.809Z [INFO] [rust:rmcp::service] task cancelled
17:27:28.809Z [INFO] [rust:rmcp::service] serve finished {"quit_reason":"Cancelled"}
Local/stdio servers appear to reconnect fast enough afterward to recover; atlassian's remote HTTP+OAuth handshake is slower and consistently loses the race, leaving it permanently failed with no automatic retry.
Confirmed not a network/auth issue independently:
- DNS resolves fine for
mcp.atlassian.com. curl https://mcp.atlassian.com/v2/mcpreturns401(server reachable, no valid header supplied in the manual test — expected).ATLASSIAN_MCP_AUTH_HEADERenv var is set.
Impact
Every /clear (a very common action) breaks the Atlassian MCP integration and requires a manual reconnect or CLI restart to restore it — significant daily friction for anyone using an HTTP-based MCP server.
Suggested fix
- Don't tear down/reinitialize already-connected MCP servers on a foreground-session handoff that isn't actually changing MCP config.
- If a reload is unavoidable, retry a server that was cancelled mid-handshake instead of leaving it permanently in
failedstate. - Respect the "MCP reload lock still held at disposal" case by waiting for the lock instead of draining the graph without it.
- Ngôn ngữ chính
- Shell
- Star
- 11.2k
- Fork
- 1.9k
- Merge trung bình
- 14 giờ 16 phút
- Pull request đã merge (30 ngày)
- 6
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của github/copilot-cli
-
triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
github/copilot-cli#4932 ·
-
triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
github/copilot-cli#4909 ·
-
triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
github/copilot-cli#4906 ·
-
triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
github/copilot-cli#4848 ·
-
area:agents area:mcp
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
github/copilot-cli#4729 ·
Tất cả issue của github/copilot-cli
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
elastic/gradle-plugins#156 ·
-
Priority/High ready-for-agent Severity/Major Type/Bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
comp/cli P3 type/docs
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
NousResearch/hermes-agent#119756 · 1 bình luận ·
-
comp: build/pipeline type: bug version: current (v17+)
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
angular/angularfire#3766 ·
-
out-of-date
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
CachyOS/CachyOS-PKGBUILDS#1903 ·