feat(supervisor): hold reconnect while a sandbox is suspended by its backend
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
調査の方向性
Start by reading the supervisor reconnect logic and the backend suspend/resume contract added in draft PR #3898; the issue does not name files or tests. The key design question is how the supervisor receives suspend and resume signals, including in the multi-tenant case. Done means suspended sandboxes trigger no connection attempts, resume restores the session without restarting the workload, and ordinary connection loss still reconnects.
索引モデルが issue の本文から書いたものです。
説明
User Story
As an operator running OpenShell sandboxes on a compute backend that suspends idle sandboxes,
I want the supervisor to hold its reconnect while a sandbox is suspended,
so that idle sandboxes stay suspended and stop using compute until they are needed.
Problem Statement
The supervisor treats any drop of its connection to the sandbox as a failure and reconnects right away. When a backend suspends a sandbox, that connection drops too. On a backend where a connection to a suspended sandbox resumes it, such as Agent Substrate, the supervisor's reconnect wakes the sandbox straight back up, so it never stays suspended.
The supervisor has no way to tell "the sandbox was suspended" from "the connection broke".
Impact / Why This Matters
Suspend and resume on a backend only saves compute if a suspended sandbox stays suspended. Today every suspend is undone by the supervisor's own reconnect.
We saw this in a prototype on Agent Substrate: OpenShell's sandbox runtime (at ec49209d) inside a Substrate microVM, with OpenShell's own supervisor reaching it through the Substrate ingress.
- Without any change, suspending the sandbox closed the connection, the supervisor redialed, and the redial resumed the sandbox shortly after the suspend.
- With a stand-in component that held the supervisor's redial until the sandbox resumed, a 15-minute suspend stayed suspended and resumed cleanly on a different machine. The runtime recovered the session after the redial, since its reconnect timeout only starts counting after resume.
The same question applies to the planned multi-tenant supervisor: if it keeps a connection open for every suspended sandbox, suspended sandboxes still get woken and still cost the supervisor a connection each.
Proposed Design
- When a backend suspends a sandbox, the supervisor learns that the sandbox is suspended, and does not reconnect to it while it stays suspended.
- When the sandbox resumes, the supervisor reconnects and the session continues, without restarting the workload.
- An unexpected connection loss, with no suspend, still triggers a reconnect as it does today.
- This works the same for today's per-sandbox supervisor and for a multi-tenant supervisor.
Draft PR #3898 (delegated backend wire contract) already adds suspend and resume to the backend contract, so the gateway knows when a sandbox is suspended. That may be the natural path for the supervisor to learn about it too. We'd defer to you on the right signal.
Acceptance Criteria
- A sandbox suspended by its backend stays suspended until it is resumed. The supervisor makes no connection attempts to it in the meantime.
- After resume, the supervisor re-establishes its session and the workload continues without a restart.
- A connection loss that is not a suspend still triggers a reconnect, as today.
- The behavior is the same with a multi-tenant supervisor, and a suspended sandbox does not hold a supervisor connection.
Alternatives Considered
- The backend refuses connections to a suspended sandbox: the redials then fail, and in our runs the supervisor exited after a refusal.
- A component next to the supervisor holds its redial while the sandbox is suspended: this is what our prototype did, and it works, but every backend would have to build one, and it has to work out state the gateway already knows.
- The runtime drops its old connection when the same supervisor redials: useful on its own, but it doesn't stop the redial from waking the sandbox.
Checklist
- I've reviewed existing issues and the published docs
- This is a design proposal, not a "please build this" request
@drew, related to the suspend and resume work in #3898. Happy to share more detail from the prototype.
- 主要言語
- Rust
- スター
- 15.4k
- フォーク
- 1.7k
- 平均マージ
- 1日 21時間
- マージ済み PR(30日)
- 358
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
NVIDIA/OpenShell のほかの issue
-
state:triage-needed
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
メンテナーはふだん 1 日以内に返信
-
state:triage-needed
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
メンテナーはふだん 1 日以内に返信
-
docs: document workspace and provider label capabilities対応中かも @johntmyers が 4 日前に担当しました。 オープンarea:docs
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
NVIDIA/OpenShell#4250 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
bug(driver-mxc): test helper fails to compile after gateway-name argument対応中かも @feloy が 5 日前に担当しました。 オープンstate:triage-needed
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 1 日以内に返信
-
bug: install.sh ignores XDG_CONFIG_HOME for the local gateway config対応中かも @fede-kamel が 9 日前に担当しました。 オープンarea:cli os:linux os:macos state:validated
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
NVIDIA/OpenShell#4042 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
NVIDIA/OpenShell の issue をすべて見る
似ている issue
-
[Bug]: Web chat input doesn't regain focus after a reply finishes対応中かも @GaijinSystems が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
zeroclaw-labs/zeroclaw#11658 ·
メンテナーはふだん 2 日以内に返信
-
good first issue help wanted
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
-
documentation
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
NuSkooler/enigma-bbs#907 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信