Sessions become unrevivable: a stale `inuse.<pid>.lock` from a crashed host is never reclaimed on open
維護者通常 1 天內回覆
還沒有人認領這個 Issue。
評估
研究方向
首先追蹤 ~/.copilot/session-state//inuse..lock 的 session-open 處理以及 copilot --server 的生命週期。重現強制終止的情況,接著檢查 Rust 工作階段執行階段中與鎖回收和主機註冊相關的部分;當已死亡或處於 zombie 狀態的 PID 不再阻止重新開啟,且工作階段無需手動移除鎖即可恢復時,主要缺陷就算完成修正。OAuth 和 events.jsonl 相關問題應作為獨立範圍處理。
由索引模型根據 Issue 內容生成。
描述
Describe the bug
Several saved sessions became impossible to reopen ("crashed out / unrevivable"). On investigation the session data was not corrupt — the append-only events.jsonl replayed cleanly and the engine logged a successful resume. The sessions were blocked by runtime/lifecycle issues, not data loss.
The primary cause is stale lock handling. Each live session writes ~/.copilot/session-state/<id>/inuse.<pid>.lock. When the owning copilot --server process exits uncleanly (crash, force-quit, OS kill), the lock file is left behind. On the next open, the session is treated as in-use / unrevivable instead of checking whether that PID is still alive.
Scanning one local store, 14 of 184 sessions (~8%) carried a fault signature; 8 held an inuse.<pid>.lock whose PID was dead. Deleting the stale lock made each session openable again, and the app then spawned a fresh host and resumed normally — confirming the data was fine and only the lock blocked recovery.
Three related lifecycle defects compound the impact (details under Additional context):
- Orphaned engine host after a failed UI attach — the engine resumes but no session host registers, leaving a live server holding the lock (which then becomes defect #1 on the next attempt).
- MCP OAuth timeouts block session readiness — resume stalls ~5 minutes on
OAuth callback timeoutbefore the session is interactive, so users abort. - Unbounded
events.jsonlmakes replay hang on large/old sessions (appears to overlap #4251).
Affected version
1.0.83 (copilot --server --stdio --no-auto-update)
Steps to reproduce the behavior
Primary defect (stale lock not reclaimed):
- Open a session, then hard-kill its
copilot --serverprocess to simulate a crash / force-quit (kill -9 <pid>). - Confirm
~/.copilot/session-state/<id>/inuse.<pid>.lockremains, now pointing at a dead PID. - Try to reopen the session → it does not revive.
rmthe stale lock file → the session reopens and resumes cleanly (a fresh host is spawned andevents.jsonlreplays without error).
Expected behavior
On open, if the lock's PID is dead or a zombie, the app should reclaim the session automatically rather than treating it as in-use. A liveness check combined with a PID start-time comparison avoids the classic PID-reuse race. Recoverable work should never be gated behind a leftover lock file that only manual filesystem surgery can clear.
Additional context
- OS: macOS · shell: zsh
Defect 2 — orphaned engine host after a failed UI attach. Engine log from a single resume: the host never registers, yet the engine reports the session resumed, so the process lingers as an orphan holding the lock:
[WARNING] [rust:copilot_runtime::session::pending_request_flow] pending-request event was not delivered to the session host
{"error":"GenericFailure, no session host is registered for session <id>"}
...
[INFO] Resumed session: <id> (logged repeatedly)
Suggested fix: tie the engine host's lifetime to successful UI host registration; if registration fails/times out, tear the host down and release the lock.
Defect 3 — MCP OAuth timeouts block resume (~5 min). With several MCP servers configured, resume stalls until interactive OAuth flows time out. MCP init begins at 20:27, first hard failures at 20:32:
[ERROR] [rust:mcp_engine::session_authorizer] OAuth login failed after returning its authorization URL
{"server_name":"...","error":"OAuth callback timeout"}
During that window the session looks hung and users abort — observed as sessions whose only recent events were resume → abort (user_initiated) → shutdown, twice in a row. Suggested fix: make MCP init lazy / non-blocking; never gate session interactivity on MCP OAuth.
Defect 4 — unbounded events.jsonl. The append-only log grows without bound; worst local session reached 38.6 MB / ~13k events with 42 compaction cycles (173 re-injected system.message events ≈ 10.7 MB). Reopening replays the whole log, so large sessions hang on open even though nothing is corrupt. This looks like the same root cause as #4251; noting here for linkage. Suggested fix: snapshot-and-truncate the event log at compaction so replay cost stays bounded.
Possibly related: #4020 (session falsely "already in use by another client"), #4098 (truncated/concatenated events on resume), #4138 (resume compaction hangs), #4251 (large-session resume OOM/hang).
All four are addressable without changing the on-disk event format. Defects 1–2 turn recoverable work into apparent data loss; 3–4 make large/older sessions feel dead on open.
- 主要語言
- Shell
- 星號
- 11.2k
- 分支
- 1.9k
- 平均合併
- 17 小時 6 分鐘
- 30 天內合併 PR
- 5
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 沒有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
github/copilot-cli 的其他 Issue
-
area:sessions
難度 2/5 1-3 小時 新手友好度 72/100
github/copilot-cli#4996 ·
維護者通常 1 天內回覆
-
triage
難度 1/5 1 小時以內 新手友好度 88/100
github/copilot-cli#4963 · 1 則留言 ·
維護者通常 1 天內回覆
-
triage
難度 2/5 1-3 小時 新手友好度 75/100
github/copilot-cli#4932 ·
維護者通常 1 天內回覆
-
triage
難度 2/5 1-3 小時 新手友好度 78/100
github/copilot-cli#4909 ·
維護者通常 1 天內回覆
-
triage
難度 2/5 1-3 小時 新手友好度 76/100
github/copilot-cli#4906 ·
維護者通常 1 天內回覆
查看 github/copilot-cli 的全部 Issue
相似的 Issue
-
area/cli
難度 2/5 1-3 小時 新手友好度 82/100
-
Remove `git-lfs`未關閉
難度 2/5 1-3 小時 新手友好度 68/100
jenkins-infra/packer-images#3114 ·
維護者通常 1 天內回覆
-
[bug] Setup fails with "Cannot find matching keyid" when an older Node's corepack is on PATH可能已有人在做 @EyalPoly 今天認領。 未關閉
難度 2/5 1-3 小時 新手友好度 72/100
維護者通常 1 天內回覆
-
status:needs-triage
難度 2/5 1-3 小時 新手友好度 88/100
PX4/PX4-Autopilot#29006 · 1 則留言 ·
維護者通常 1 天內回覆
-
難度 2/5 1-3 小時 新手友好度 76/100
idean3885/claude-ops-agent#633 ·
維護者通常 1 天內回覆