Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

feat(supervisor): hold reconnect while a sandbox is suspended by its backend

オープン
#4,352 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
35/100
issue の種類
機能追加
明瞭さ
おおむね明確
活発さ
活発
技術スタック
rust
領域
backend

調査の方向性

Start by reading the supervisor reconnect logic and the backend suspend/resume contract added in draft PR #3898; the issue does not name files or tests. The key design question is how the supervisor receives suspend and resume signals, including in the multi-tenant case. Done means suspended sandboxes trigger no connection attempts, resume restores the session without restarting the workload, and ordinary connection loss still reconnects.

索引モデルが issue の本文から書いたものです。

説明

state:triage-needed
User Story

As an operator running OpenShell sandboxes on a compute backend that suspends idle sandboxes,
I want the supervisor to hold its reconnect while a sandbox is suspended,
so that idle sandboxes stay suspended and stop using compute until they are needed.

Problem Statement

The supervisor treats any drop of its connection to the sandbox as a failure and reconnects right away. When a backend suspends a sandbox, that connection drops too. On a backend where a connection to a suspended sandbox resumes it, such as Agent Substrate, the supervisor's reconnect wakes the sandbox straight back up, so it never stays suspended.

The supervisor has no way to tell "the sandbox was suspended" from "the connection broke".

Impact / Why This Matters

Suspend and resume on a backend only saves compute if a suspended sandbox stays suspended. Today every suspend is undone by the supervisor's own reconnect.

We saw this in a prototype on Agent Substrate: OpenShell's sandbox runtime (at ec49209d) inside a Substrate microVM, with OpenShell's own supervisor reaching it through the Substrate ingress.

  • Without any change, suspending the sandbox closed the connection, the supervisor redialed, and the redial resumed the sandbox shortly after the suspend.
  • With a stand-in component that held the supervisor's redial until the sandbox resumed, a 15-minute suspend stayed suspended and resumed cleanly on a different machine. The runtime recovered the session after the redial, since its reconnect timeout only starts counting after resume.

The same question applies to the planned multi-tenant supervisor: if it keeps a connection open for every suspended sandbox, suspended sandboxes still get woken and still cost the supervisor a connection each.

Proposed Design
  • When a backend suspends a sandbox, the supervisor learns that the sandbox is suspended, and does not reconnect to it while it stays suspended.
  • When the sandbox resumes, the supervisor reconnects and the session continues, without restarting the workload.
  • An unexpected connection loss, with no suspend, still triggers a reconnect as it does today.
  • This works the same for today's per-sandbox supervisor and for a multi-tenant supervisor.

Draft PR #3898 (delegated backend wire contract) already adds suspend and resume to the backend contract, so the gateway knows when a sandbox is suspended. That may be the natural path for the supervisor to learn about it too. We'd defer to you on the right signal.

Acceptance Criteria
  • A sandbox suspended by its backend stays suspended until it is resumed. The supervisor makes no connection attempts to it in the meantime.
  • After resume, the supervisor re-establishes its session and the workload continues without a restart.
  • A connection loss that is not a suspend still triggers a reconnect, as today.
  • The behavior is the same with a multi-tenant supervisor, and a suspended sandbox does not hold a supervisor connection.
Alternatives Considered
  • The backend refuses connections to a suspended sandbox: the redials then fail, and in our runs the supervisor exited after a refusal.
  • A component next to the supervisor holds its redial while the sandbox is suspended: this is what our prototype did, and it works, but every backend would have to build one, and it has to work out state the gateway already knows.
  • The runtime drops its old connection when the same supervisor redials: useful on its own, but it doesn't stop the redial from waking the sandbox.
Checklist
  • I've reviewed existing issues and the published docs
  • This is a design proposal, not a "please build this" request

@drew, related to the suspend and resume work in #3898. Happy to share more detail from the prototype.

主要言語
Rust
スター
15.4k
フォーク
1.7k
平均マージ
1日 21時間
マージ済み PR(30日)
358

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/OpenShell のほかの issue

NVIDIA/OpenShell の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。