Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

Session wedges permanently when a queued-lane message lands at turn end (idle finalization suppressed, queue never drains)

未關閉
#4,755 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

評估

難度
4/5
預估耗時
3-5 天
新手友好度
45/100
Issue 類型
缺陷
描述清晰度
基本清楚
活躍度
活躍
技術堆疊
rust

研究方向

從發出 Session idle finalization suppressed 的 copilot_runtime::session::registry 路徑開始,接著追蹤 queued-lane drain 和 send_session_message 的進入點。使用 delivery_mode: \"enqueue\",透過長時間執行的父工作階段和子工作階段重現問題;完成標準是:排隊中的訊息不能讓工作階段同時處於非 idle 且未 drain 的狀態,且即時訊息仍然可傳送。

由索引模型根據 Issue 內容生成。

描述

area:agents area:sessions
Describe the bug

A session can end a turn and enter a permanently wedged state: neither idle nor
running. It accepts no further input, the app shows it as stopped, and queuing a
message into it silently does nothing. The only recovery is to kill the process.

The process does not crash. It stays alive, Responding=True, at zero CPU, with
no ERROR or WARN line anywhere in its log, and its last session events are a
clean assistant.turn_end. It is parked, not looping and not faulted.

The signature is a single log line, emitted 91 ms after the final turn end:

2026-09-07T08:44:37.505Z  assistant.turn_end                                (events.jsonl)
2026-09-07T08:44:37.596Z [DEBUG] [rust:copilot_runtime::session::registry]
  Session idle finalization suppressed
  {"queue_processing":false,"queued_lane":true,"immediate_lane":false,"host_pending_send":false}

A message arrived on the queued lane while the turn was running. At turn end
the registry correctly declined to finalize the session as idle because the lane
was non-empty, but the queue processor was never started (queue_processing: false) and nothing ever starts it. So the session cannot go idle (the queue is
non-empty) and cannot drain the queue (nothing is processing it). A lost wakeup
between turn-end idle finalization and the queued-lane drain.

The immediate lane does not rescue it. After the wedge I sent the session a
send_session_message with delivery_mode: "immediate". It was recorded as a
pending_messages_modified telemetry event and then never delivered, never
written to events.jsonl, and never processed. Once wedged, both lanes are dead.

Which sessions hit it. I grepped all 62 process logs on this machine for the
suppression line. It appears in exactly three, and all three are the same
workload shape: a long-running parent session that spawns several child sessions
with coordinate_with_creator: true and then runs multi-minute turns, so it
takes inbound cross-session messages at unpredictable moments. The examined
session made 12 create_session calls and received 11 inbound cross-session
messages before wedging.

The lane is the discriminator. In that same session, three earlier inbound
cross-session messages arrived as delivery="steering" (the immediate lane, at
08:32:00Z, 08:33:46Z and 08:34:30Z) and were all consumed normally mid-turn. The
message that wedged the session went to the queued lane. Since
send_session_message defaults to delivery_mode: "enqueue", any child session
replying to its parent with the default lands on the lane that can wedge.

That makes this reachable by ordinary multi-session orchestration, which is a
documented feature, not an exotic configuration.

Affected version

1.0.83-5 (also observed on the immediately preceding build; three occurrences
across 2026-09-06 and 2026-09-07). Running inside the GitHub Copilot desktop app,
which drives this CLI as one process per session.

Steps to reproduce the behavior

It is a race, so it reproduces probabilistically rather than deterministically.
Three occurrences in two days on a machine running this pattern continuously.

  1. Start a parent session and have it spawn several child sessions with
    create_session, passing coordinate_with_creator: true.
  2. Keep the parent busy in long turns (multiple minutes each).
  3. Have the children report back with send_session_message using the default
    delivery_mode (enqueue), at times they choose, so at least one message
    lands close to the instant a parent turn ends.
  4. Occasionally the parent's final turn end logs Session idle finalization suppressed with queued_lane: true and queue_processing: false, and the
    session is wedged from that moment on.

To confirm a wedge rather than a long turn: the process is alive and responding,
CPU is flat at zero, the log has no error, and no further events are appended to
events.jsonl.

Expected behavior

Suppressing idle finalization because the queued lane is non-empty should
guarantee the queue processor is subsequently started. The two decisions want to
be one atomic transition, so a session cannot end up in a state where it is
ineligible for idle and has nothing scheduled to drain the queue.

Failing that, either of these would make it self-healing rather than terminal:

  • a watchdog that re-checks a session which is non-idle with queue_processing: false and no in-flight turn, and starts the drain;
  • an immediate-lane delivery forcing a drain attempt, so a wedged session can
    be recovered by messaging it instead of by killing the process.

It would also help if pending_messages_modified were not reported as success to
the sender when the target session can never consume the message. From the
caller's side the send appears to succeed.

Additional context

Recovery, for anyone else who hits this: find the pid from the
inuse.<pid>.lock file in ~/.copilot/session-state/<sessionId>/ and stop that
process. The app restarts the session, replays a "Continue from where you left
off" prompt, and resumes normally. events.jsonl history and the git worktree
are untouched; only the undelivered pending messages are lost, and they were
never going to be delivered. I confirmed this: the wedged session restarted and
took a new turn 3 minutes later.

Diagnosis is straightforward if you have the logs, since the suppression line
names the exact state:

Get-ChildItem ~/.copilot/logs/*.log |
  ForEach-Object { Select-String -Path $_ -Pattern 'Session idle finalization suppressed' }

Environment

  • Operating system: Windows 11 Enterprise
  • CPU architecture: AMD64 (x86_64)
  • Host: GitHub Copilot desktop app (one CLI process per session), sessions are
    git worktree-backed
  • Shell: PowerShell

Observed impact. The wedged session was coordinating a multi-session
workstream, so the failure is not merely a lost message: the parent stops
supervising children that are still running, and there is no signal to the user
beyond the session appearing stopped. I have had to work around it by changing
our agent definitions to send all child-to-parent reports with delivery_mode: "immediate" and to stop setting notify_on_idle, purely to keep traffic off the
queued lane. That reduces the exposure but obviously does not close the race.

主要語言
Shell
星號
11.2k
分支
1.9k
平均合併
17 小時 6 分鐘
30 天內合併 PR
5

環境準備

  • 沒有 Dockerfile 或 Docker Compose 檔案
  • 沒有 Pull Request 範本
  • 閱讀貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

github/copilot-cli 的其他 Issue

查看 github/copilot-cli 的全部 Issue

相似的 Issue

更多 Shell/Bash Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。