bug(kubernetes): sandbox stuck at `Provisioning` with `replicas: 0` has no recovery path (`start`/`stop` both refuse)
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 30/100
Hướng nghiên cứu
Start with the phase gate in start_sandbox in crates/openshell-server/src/compute/mod.rs (~L1784) and the reconcile fallthrough near derive_phase (~L6863), then see how crates/openshell-driver-kubernetes/src/driver.rs suspend_sandbox_runtime_after_dependency_failure (~L4111) sets replicas to 0. Done means a sandbox stuck in Provisioning with an empty backend either gets corrected by reconcile or can be started through the idempotent-restart path, with the workspace PVC kept. The issue offers two directions, so agree the approach with maintainers before coding.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
A sandbox can end up with the gateway's Sandbox record at phase Provisioning
while the backend Sandbox CR is at spec.replicas: 0 with no pod. In that state
there is no supported recovery path:
sandbox startis rejected — the phase gate only admitsStopped,
Completed,Starting, a failed-main-processError, or a provisioning timeout.sandbox stopis rejected — it requiresReady.- The reconcile sweep does not correct the record: with
replicas: 0/ no pod, the
Provisioningcase falls through to the reconcile/derive_phasedefault and the
stale value is perpetuated rather than corrected.
The only exits are delete + recreate, or out-of-band edits to the gateway's store.
This is a distinct variant of the "unrecoverable phase" bugs #3308
and #4209, which are Error-phase. Here the record is
Provisioning, so the recovery work in #4231 (allow stop from
Error) does not reach it, and start still rejects it. The closest mechanism match is
#2932 (a stale condition perpetuated by reconcile), fixed by
#2933.
Environment
- OpenShell:
0.1.2(gatewayghcr.io/nvidia/openshell/gateway:0.1.2, supervisor0.1.2, CLI0.1.2) - Driver: Kubernetes (agent-sandbox
v0.4.6, CRDsandboxes.agents.x-k8s.iov1alpha1) - Kubernetes:
v1.27.3(kind)
Reproduction
Non-deterministic (a race). Observed triggers:
- A node/machine reboot while the gateway is bootstrapping a sandbox — the gateway pod
restarts during cluster stabilization, and a transient dependency failure fires
suspend_sandbox_runtime_after_dependency_failure, resettingreplicas: 0while the
record is stillProvisioning. - A
stopimmediately followed bystart, leaving the record mid-transition.
Observed state
$ openshell sandbox list
NAME ... Provisioning
$ kubectl get sandboxes -n <ns> -o yaml
# spec.replicas: 0
# no openshell.ai/sandbox-runtime-bootstrapping annotation
# (no sandbox pod in the namespace)
$ openshell sandbox start <name>
Error: sandbox must be Stopped, Completed, or a failed main-process Error to start (current phase: Provisioning)
$ openshell sandbox stop <name>
Error: sandbox must be Ready to stop (current phase: Provisioning)
Root cause (pointers; as of e1f3c82ca)
- The record's
status.phaseisProvisioningbut the backend CR is empty.
start_sandboxrejects it because the gate only admits
Stopped | Completed | Starting | is_failed_main_process_result | provisioning_timeout
(crates/openshell-server/src/compute/mod.rs,start_sandbox, ~L1784). - The reconcile sweep re-derives phase from the driver snapshot; with no pod /
replicas: 0, theProvisioningcase falls through to the_ => phasedefault
(~L6564, fed byderive_phase~L6863), so the stale value is kept. suspend_sandbox_runtime_after_dependency_failure
(crates/openshell-driver-kubernetes/src/driver.rs, ~L4111) is what drives
replicas: 0on the dependency-failure path.
Expected
One of:
- The reconcile loop corrects a
Provisioningrecord whose backend reports the
sandbox as empty (no pod /replicas: 0) instead of perpetuating it — cf. #2933, which
madederive_phaseletReadywin over a staleSuspended; or startaccepts a staleProvisioningwhose backend is empty and routes it through the
existing idempotent-restart path — reusing the persisted identity and re-attaching the
existing workspace PVC — the same way theStartingbranch already does (~L1800).
Actual
No recovery short of delete + recreate. An out-of-band edit of the gateway store
(flipping Sandbox.status.phase from Provisioning to Stopped) does make start
succeed and preserves the workspace — confirming this is a fixable record desync rather
than data loss — but it is not a supported operator path.
Proposed direction
- Extend #4231's recovery approach to the
Provisioning-with-empty-backend case (not
justError), or - correct the stale
Provisioninginderive_phase/reconcile (precedent: #2933).
Either way, start from this state should reuse the persisted identity and the existing
workspace PVC (idempotent restart), not require a delete.
References
- #3308 — unrecoverable
Error, stop/start refuse (same symptom class) - #4209 — gateway restart → unrecoverable
Error(same trigger family) - #4231 — open: allow stopping errored sandboxes for recovery (in-flight recovery mechanism)
- #2932 / #2933 — stale condition perpetuated by reconcile; fix in
derive_phase(closest mechanism) - #4078 — serialize lifecycle cleanup with sandbox restart (related stop/start race)
- Ngôn ngữ chính
- Rust
- Star
- 15.4k
- Fork
- 1.7k
- Merge trung bình
- 1 ngày 21 giờ
- Pull request đã merge (30 ngày)
- 358
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/OpenShell
-
state:triage-needed
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
Maintainer thường phản hồi trong vòng 1 ngày
-
state:triage-needed
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
Maintainer thường phản hồi trong vòng 1 ngày
-
docs: document workspace and provider label capabilitiesCó thể đã có người làm @johntmyers đã nhận 4 ngày trước. Đang mởarea:docs
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
NVIDIA/OpenShell#4250 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug(driver-mxc): test helper fails to compile after gateway-name argumentCó thể đã có người làm @feloy đã nhận 5 ngày trước. Đang mởstate:triage-needed
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 1 ngày
-
bug: install.sh ignores XDG_CONFIG_HOME for the local gateway configCó thể đã có người làm @fede-kamel đã nhận 9 ngày trước. Đang mởarea:cli os:linux os:macos state:validated
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
NVIDIA/OpenShell#4042 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của NVIDIA/OpenShell
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
NuSkooler/enigma-bbs#907 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
EasyTier/EasyTier#2672 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
XmlFragment children, successors and siblings stop at the first child that is not an XML typeĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
bug good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
repowise-dev/repowise#3374 ·
Maintainer thường phản hồi trong vòng 1 ngày