Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

bug: Kubernetes v1beta1 stop-start can hang waiting for supervisor readiness

未关闭
#3,566 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
48/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
github-actions, kubernetes

调研方向

Start by running the Kubernetes E2E workflow with the sandbox-lifecycle stop/start scenario and inspect the replacement supervisor pod and readiness-probe logs. Trace the replacement supervisor's startup and gateway reconnect path; done means the v1beta1 scenario completes reliably, a regression or stress test covers it, and failures identify the blocked startup stage rather than only a 120-second CLI timeout.

由索引模型根据 Issue 内容生成。

描述

topic:testing

User Story

As an OpenShell contributor, I want the Kubernetes sandbox stop/start lifecycle test to complete reliably, so that required CI reflects product regressions instead of intermittent supervisor startup hangs.

Problem Statement

The Kubernetes Agent Sandbox v1beta1 E2E lane intermittently times out during the sandbox-lifecycle stop/start scenario. The gateway completes StartSandbox successfully and Kubernetes starts the replacement supervisor pod, but the supervisor never creates its readiness socket or reconnects to the gateway. Its only log line is Starting sandbox supervision.

The same v1beta1 lane passed on the preceding commit with identical Kubernetes code, while the v1alpha1, workspace-managed, workspace-operator, and external-driver Kubernetes lanes passed in the failing run.

Failed job: https://github.com/NVIDIA/OpenShell/actions/runs/35776641970/job/106915241802

Manual retry: https://github.com/NVIDIA/OpenShell/actions/runs/35776641970/job/106938509231

Impact / Why This Matters

This intermittently fails the required Core E2E gate, delays otherwise valid pull requests, and consumes additional CI capacity through manual reruns. A generic 120-second CLI timeout also makes the underlying supervisor startup stage difficult to diagnose. Retrying the job is the current workaround, but it does not prevent recurrence or protect the stop/start lifecycle from real regressions.

Acceptance Criteria

  • The v1beta1 stop/start scenario no longer intermittently stalls after the replacement supervisor pod starts.
  • A regression test or repeatable stress test covers stop followed by start and verifies that the replacement supervisor reconnects and becomes ready.
  • If supervisor startup cannot complete, CI reports the blocked startup stage or cause instead of only timing out at the CLI after 120 seconds.
  • The existing v1alpha1 and other Kubernetes deployment-mode E2E lanes continue to pass.

Reproduction Steps

  1. Run the Kubernetes E2E workflow with Agent Sandbox v0.5.0 and the sandbox-lifecycle conformance scenario.
  2. Create a sandbox and wait for it to become ready.
  3. Stop the sandbox.
  4. Start the sandbox again.
  5. Intermittently, observe openshell sandbox start time out after 120 seconds while the replacement supervisor pod remains unready.

Environment

  • OpenShell: 0c6c60c1eb70b1c4bc8e8a4f6e06d74ea09069bc
  • OS: Linux amd64 GitHub Actions runner
  • Runtime, deployment, or integration: kind Kubernetes cluster, Agent Sandbox API v1beta1, Agent Sandbox controller v0.5.0

Logs

[sandbox-lifecycle/stop-start/start] timed out after 120.0s
command: openshell sandbox start ct-869gjwhrud-ss

[pod/os-supervisor-.../supervisor] Starting sandbox supervision
Readiness probe failed: connect supervisor readiness socket /run/openshell/health.sock
主要语言
Rust
星标
8.7k
派生
1.3k
平均合并
2 天 6 小时
30 天内合并 PR
301

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

NVIDIA/OpenShell 的其他 Issue

查看 NVIDIA/OpenShell 的全部 Issue

相似的 Issue

更多 Rust Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。