bug: local sandbox sessions fail to recover after host sleep
@drew 已经在做这个了。
开始于 2026年9月23日。
评估
这个 Issue 还没有评估数据。
描述
User Story
As a developer running OpenShell locally, I want an attached sandbox to survive laptop sleep and wake, so that I can resume the same canonical process and workspace without recreating the sandbox.
Problem Statement
Local Docker, Podman, and VM gateways use non-expiring bootstrap credentials when gateway_jwt.ttl_secs is omitted, but launch-scoped gateway and Sandbox Protocol credentials currently fall back to a 15-minute lifetime. A laptop can remain suspended beyond that lifetime without giving the supervisor an opportunity to refresh. After wake, the sandbox transport can be closed as expired and sandbox connect does not recover an established SSH transport interruption.
This creates inconsistent local behavior: start, stop, and exec establish fresh command paths, while an interactive connect session can exit or fail to reattach after sleep.
Impact / Why This Matters
Developers lose long-running interactive sessions merely by closing a laptop. The practical workaround is to rerun commands, restart components, or recreate the sandbox, which can discard process state and interrupts the expected persistent-sandbox workflow. Increasing a finite TTL only changes how long the laptop may sleep before failure and does not make local sessions robust.
Acceptance Criteria
- Omitting
gateway_jwt.ttl_secson local Docker, Podman, and VM gateways produces non-expiring launch-scoped gateway credentials. - The same omission produces non-expiring Sandbox Protocol credentials and propagates that state through refresh responses and clients.
- Non-expiring credentials do not schedule an immediate or finite sandbox connection deadline.
- Shared deployments can continue to configure positive credential TTLs, and Kubernetes retains its positive default.
-
sandbox connectretries an established SSH transport failure for a bounded period and reattaches to the same canonical main process. - Initial authentication failures, sandbox lifecycle failures, and clean process exits are not retried.
- Automated tests cover non-expiring credential propagation, gateway restart or stop-start behavior, and forced SSH transport recovery.
Reproduction Steps
- Start a local Docker, Podman, or VM gateway with
gateway_jwt.ttl_secsomitted. - Create a persistent sandbox with a long-running canonical main process.
- Attach with
openshell sandbox connect. - Suspend the host laptop for longer than 15 minutes.
- Wake the laptop and attempt to continue or reconnect to the same sandbox.
- Observe that the interactive connection exits or cannot resume even though lifecycle commands may still establish fresh command paths.
Environment
- OpenShell: main before PR #3573
- OS: laptop host with suspend and resume
- Runtime, deployment, or integration: local Docker, Podman, or VM gateway
Proposed Fix
Propagate the local non-expiring JWT configuration to both launch-scoped credential profiles and supervise the SSH child used by sandbox connect, allowing bounded transport recovery and reattachment to the same canonical process.
Implementation: #3573
- 主要语言
- Rust
- 星标
- 8.7k
- 派生
- 1.3k
- 平均合并
- 2 天 6 小时
- 30 天内合并 PR
- 301
贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/OpenShell 的其他 Issue
-
area:docs
难度 1/5 1 小时以内 新手友好度 88/100
-
state:triage-needed
难度 2/5 1-3 小时 新手友好度 82/100
-
area:cli state:validated
难度 2/5 1-3 小时 新手友好度 72/100
-
state:triage-needed
难度 1/5 1 小时以内 新手友好度 90/100
-
area:build spike state:review-ready state:stale
难度 2/5 半天 新手友好度 68/100
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 75/100
TheLarkInn/aipm#2413 ·
-
documentation
难度 1/5 1 小时以内 新手友好度 90/100
alexgorbatchev/simple-ptt#15 ·
-
tooling
难度 2/5 1-3 小时 新手友好度 75/100
-
todo:ticket
难度 2/5 1-3 小时 新手友好度 70/100
-
难度 2/5 1-3 小时 新手友好度 75/100
taikoxyz/taiko-mono#22168 · 1 条评论 ·