bug: Kubernetes supervisor Pod is BestEffort QoS, so eviction terminates the sandbox
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 45/100
- Issue 类型
- 缺陷
- 描述清晰度
- 基本清楚
- 活跃度
- 活跃
- 技术栈
- kubernetes, rust
调研方向
先从 crates/openshell-driver-kubernetes/src/sandbox_runtime.rs 中的 supervisor_pod() 和 KubernetesSandboxRuntimeConfig 开始,然后阅读 crates/openshell-sandbox-backend/src/sandbox_auth.rs 中的 SandboxConnectionRegistry。检查 docs/reference/gateway-config.mdx 和现有的 Kubernetes 驱动测试。完成的标准是:默认和可配置的 supervisor 资源已渲染并记录在文档中,资源测试通过,并且已明确决定和记录 session-rotation 行为。
由索引模型根据 Issue 内容生成。
描述
User Story
As a platform operator running OpenShell sandboxes on a shared Kubernetes cluster, I want the per-sandbox supervisor Pod to declare resource requests, so that node memory pressure does not silently terminate running agent sandboxes.
Problem Statement
The RFC 0012 Kubernetes topology creates a companion supervisor Pod per sandbox. That Pod is built in supervisor_pod() (crates/openshell-driver-kubernetes/src/sandbox_runtime.rs) with no resources block, and the driver config exposes no field to add one — KubernetesSandboxRuntimeConfig offers only supervisor_image and supervisor_image_pull_policy.
With no requests or limits, the supervisor Pod is assigned BestEffort QoS, which makes it the first candidate the kubelet evicts under node memory pressure. The workload Pod does receive resources (template limits are mirrored into requests), so the pair is asymmetric: the more carefully a sandbox is sized, the larger the gap between the two Pods' eviction priority.
Eviction of the supervisor is not recoverable in place:
- The supervisor Pod sets
restartPolicy: Never, so the kubelet does not restart it. - On losing the boundary connection the sandbox freezes the owned workload process tree and opens a 30-second reconnect window.
- That window admits only the same supervisor process.
SandboxConnectionRegistry::attach(crates/openshell-sandbox-backend/src/sandbox_auth.rs) pins the firstsupervisor_instance_idit accepts and returnsWrongSupervisorInstancefor any other. A recreated Pod is a new process with a new ephemeral instance ID. - The session-rotation lineage that would admit a signed successor is not wired up on this path —
SandboxConnectionRegistry::newaccepts_session_idand_session_rotationand ignores both.
When the window expires, the sandbox terminates the workload. An unrelated memory-hungry Pod scheduled onto the same node can therefore destroy running agent sandboxes.
Impact / Why This Matters
Agent sandboxes are long-running and carry uncommitted work in the PVC-backed /sandbox workspace. Losing one mid-run costs the user that session, and the trigger is a node condition the sandbox owner has no visibility into or control over. Because the supervisor is also the component holding the gateway session, the failure surfaces as a sandbox that stops responding rather than as an obvious infrastructure eviction.
The only workaround today is cluster-side: apply a LimitRange to the sandbox namespace so Kubernetes injects default requests into the supervisor Pod. That is insufficient because it is invisible to OpenShell, applies the same defaults to every Pod in the namespace including workload Pods that already carry their own sizing, and must be rediscovered and reapplied by every operator independently. It also cannot be tuned per sandbox profile — a supervisor doing TLS interception and L7 inspection for a busy agent has a materially different footprint from one relaying an idle session.
A related consequence worth deciding on separately: because the scheduler sees zero requests for supervisor Pods, node capacity planning undercounts them entirely. On a node packed with sandboxes this makes memory pressure more likely, which feeds back into the eviction path above.
Acceptance Criteria
- The Kubernetes supervisor Pod declares CPU and memory requests by default, placing it above BestEffort QoS.
- Operators can override the supervisor Pod's requests and limits through
[openshell.drivers.kubernetes]configuration. - The defaults are documented in
docs/reference/gateway-config.mdx. - A unit test asserts the rendered supervisor PodSpec carries resource requests.
- Decide and record whether an evicted or otherwise replaced supervisor should be able to resume a running generation through the existing session-rotation claims, or whether termination remains the intended behavior.
Reproduction Steps
- Deploy the gateway with the Kubernetes compute driver and create a sandbox.
- Inspect the supervisor Pod's QoS class:
Observekubectl get pod os-supervisor-<sandbox-id> -n <sandbox-namespace> \ -o jsonpath='{.status.qosClass}{"\n"}'BestEffort. - Drive the node hosting that Pod into memory pressure, or evict it directly:
kubectl delete pod os-supervisor-<sandbox-id> -n <sandbox-namespace> - Observe that the supervisor Pod is not restarted, and that the sandbox's workload is terminated after the reconnect window elapses rather than recovering.
Environment
- OpenShell:
mainat c1f2e718 (RFC 0012 isolation architecture, PR #2942) - Compute driver: Kubernetes
- Applies to all
workspace_modevalues
- 主要语言
- Rust
- 星标
- 13.2k
- 派生
- 1.6k
- 平均合并
- 1 天 20 小时
- 30 天内合并 PR
- 333
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
NVIDIA/OpenShell 的其他 Issue
-
bug: install.sh ignores XDG_CONFIG_HOME for the local gateway config可能已有人在做 @fede-kamel 于 2 天前认领。 未关闭area:cli os:linux os:macos state:validated
难度 2/5 1-3 小时 新手友好度 88/100
NVIDIA/OpenShell#4042 · 2 条评论 ·
维护者通常 1 天内回复
-
state:triage-needed
难度 2/5 1-3 小时 新手友好度 72/100
NVIDIA/OpenShell#3995 · 2 条评论 ·
维护者通常 1 天内回复
-
state:triage-needed
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
state:triage-needed
难度 1/5 1-3 小时 新手友好度 88/100
维护者通常 1 天内回复
-
bug: Kubernetes setup guide points at a nonexistent agent-sandbox manifest.yaml asset可能重新可做 关联的 PR 已关闭且未合并。 未关闭area:docs
难度 1/5 1 小时以内 新手友好度 88/100
维护者通常 1 天内回复
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 92/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 72/100
rust-windowing/winit#4731 ·
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 78/100
维护者通常 1 天内回复
-
area:cli bug priority:high
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 1 天内回复
-
component:sight
难度 2/5 1-3 小时 新手友好度 78/100
agentic-os-org/ANOLISA#4622 · 1 条评论 ·
维护者通常 1 天内回复