Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

bug: Kubernetes supervisor Pod is BestEffort QoS, so eviction terminates the sandbox

オープン
#3,415 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
5/5
見積もり時間
1週間以上
初心者へのやさしさ
45/100
issue の種類
バグ
明瞭さ
おおむね明確
活発さ
活発
技術スタック
kubernetes, rust

調査の方向性

crates/openshell-driver-kubernetes/src/sandbox_runtime.rs の supervisor_pod() と KubernetesSandboxRuntimeConfig から始め、次に crates/openshell-sandbox-backend/src/sandbox_auth.rs の SandboxConnectionRegistry を読みます。docs/reference/gateway-config.mdx と既存の Kubernetes ドライバーテストを確認します。supervisor のデフォルトリソースと設定可能なリソースがレンダリングおよび文書化され、リソーステストがパスし、セッションローテーションの動作が明示的に決定されて記録されていれば完了です。

索引モデルが issue の本文から書いたものです。

説明

state:triage-needed

User Story

As a platform operator running OpenShell sandboxes on a shared Kubernetes cluster, I want the per-sandbox supervisor Pod to declare resource requests, so that node memory pressure does not silently terminate running agent sandboxes.

Problem Statement

The RFC 0012 Kubernetes topology creates a companion supervisor Pod per sandbox. That Pod is built in supervisor_pod() (crates/openshell-driver-kubernetes/src/sandbox_runtime.rs) with no resources block, and the driver config exposes no field to add one — KubernetesSandboxRuntimeConfig offers only supervisor_image and supervisor_image_pull_policy.

With no requests or limits, the supervisor Pod is assigned BestEffort QoS, which makes it the first candidate the kubelet evicts under node memory pressure. The workload Pod does receive resources (template limits are mirrored into requests), so the pair is asymmetric: the more carefully a sandbox is sized, the larger the gap between the two Pods' eviction priority.

Eviction of the supervisor is not recoverable in place:

  1. The supervisor Pod sets restartPolicy: Never, so the kubelet does not restart it.
  2. On losing the boundary connection the sandbox freezes the owned workload process tree and opens a 30-second reconnect window.
  3. That window admits only the same supervisor process. SandboxConnectionRegistry::attach (crates/openshell-sandbox-backend/src/sandbox_auth.rs) pins the first supervisor_instance_id it accepts and returns WrongSupervisorInstance for any other. A recreated Pod is a new process with a new ephemeral instance ID.
  4. The session-rotation lineage that would admit a signed successor is not wired up on this path — SandboxConnectionRegistry::new accepts _session_id and _session_rotation and ignores both.

When the window expires, the sandbox terminates the workload. An unrelated memory-hungry Pod scheduled onto the same node can therefore destroy running agent sandboxes.

Impact / Why This Matters

Agent sandboxes are long-running and carry uncommitted work in the PVC-backed /sandbox workspace. Losing one mid-run costs the user that session, and the trigger is a node condition the sandbox owner has no visibility into or control over. Because the supervisor is also the component holding the gateway session, the failure surfaces as a sandbox that stops responding rather than as an obvious infrastructure eviction.

The only workaround today is cluster-side: apply a LimitRange to the sandbox namespace so Kubernetes injects default requests into the supervisor Pod. That is insufficient because it is invisible to OpenShell, applies the same defaults to every Pod in the namespace including workload Pods that already carry their own sizing, and must be rediscovered and reapplied by every operator independently. It also cannot be tuned per sandbox profile — a supervisor doing TLS interception and L7 inspection for a busy agent has a materially different footprint from one relaying an idle session.

A related consequence worth deciding on separately: because the scheduler sees zero requests for supervisor Pods, node capacity planning undercounts them entirely. On a node packed with sandboxes this makes memory pressure more likely, which feeds back into the eviction path above.

Acceptance Criteria

  • The Kubernetes supervisor Pod declares CPU and memory requests by default, placing it above BestEffort QoS.
  • Operators can override the supervisor Pod's requests and limits through [openshell.drivers.kubernetes] configuration.
  • The defaults are documented in docs/reference/gateway-config.mdx.
  • A unit test asserts the rendered supervisor PodSpec carries resource requests.
  • Decide and record whether an evicted or otherwise replaced supervisor should be able to resume a running generation through the existing session-rotation claims, or whether termination remains the intended behavior.

Reproduction Steps

  1. Deploy the gateway with the Kubernetes compute driver and create a sandbox.
  2. Inspect the supervisor Pod's QoS class:
    kubectl get pod os-supervisor-<sandbox-id> -n <sandbox-namespace> \
      -o jsonpath='{.status.qosClass}{"\n"}'
    
    Observe BestEffort.
  3. Drive the node hosting that Pod into memory pressure, or evict it directly:
    kubectl delete pod os-supervisor-<sandbox-id> -n <sandbox-namespace>
    
  4. Observe that the supervisor Pod is not restarted, and that the sandbox's workload is terminated after the reconnect window elapses rather than recovering.

Environment

  • OpenShell: main at c1f2e718 (RFC 0012 isolation architecture, PR #2942)
  • Compute driver: Kubernetes
  • Applies to all workspace_mode values
主要言語
Rust
スター
8.7k
フォーク
1.3k
平均マージ
2日 6時間
マージ済み PR(30日)
297

コントリビューションガイド

コントリビューションガイドを開く

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

NVIDIA/OpenShell のほかの issue

NVIDIA/OpenShell の issue をすべて見る

似ている issue

Rust の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。