bug(kubernetes): sandbox stuck at `Provisioning` with `replicas: 0` has no recovery path (`start`/`stop` both refuse)
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 30/100
Direzione di ricerca
Start with the phase gate in start_sandbox in crates/openshell-server/src/compute/mod.rs (~L1784) and the reconcile fallthrough near derive_phase (~L6863), then see how crates/openshell-driver-kubernetes/src/driver.rs suspend_sandbox_runtime_after_dependency_failure (~L4111) sets replicas to 0. Done means a sandbox stuck in Provisioning with an empty backend either gets corrected by reconcile or can be started through the idempotent-restart path, with the workspace PVC kept. The issue offers two directions, so agree the approach with maintainers before coding.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
A sandbox can end up with the gateway's Sandbox record at phase Provisioning
while the backend Sandbox CR is at spec.replicas: 0 with no pod. In that state
there is no supported recovery path:
sandbox startis rejected — the phase gate only admitsStopped,
Completed,Starting, a failed-main-processError, or a provisioning timeout.sandbox stopis rejected — it requiresReady.- The reconcile sweep does not correct the record: with
replicas: 0/ no pod, the
Provisioningcase falls through to the reconcile/derive_phasedefault and the
stale value is perpetuated rather than corrected.
The only exits are delete + recreate, or out-of-band edits to the gateway's store.
This is a distinct variant of the "unrecoverable phase" bugs #3308
and #4209, which are Error-phase. Here the record is
Provisioning, so the recovery work in #4231 (allow stop from
Error) does not reach it, and start still rejects it. The closest mechanism match is
#2932 (a stale condition perpetuated by reconcile), fixed by
#2933.
Environment
- OpenShell:
0.1.2(gatewayghcr.io/nvidia/openshell/gateway:0.1.2, supervisor0.1.2, CLI0.1.2) - Driver: Kubernetes (agent-sandbox
v0.4.6, CRDsandboxes.agents.x-k8s.iov1alpha1) - Kubernetes:
v1.27.3(kind)
Reproduction
Non-deterministic (a race). Observed triggers:
- A node/machine reboot while the gateway is bootstrapping a sandbox — the gateway pod
restarts during cluster stabilization, and a transient dependency failure fires
suspend_sandbox_runtime_after_dependency_failure, resettingreplicas: 0while the
record is stillProvisioning. - A
stopimmediately followed bystart, leaving the record mid-transition.
Observed state
$ openshell sandbox list
NAME ... Provisioning
$ kubectl get sandboxes -n <ns> -o yaml
# spec.replicas: 0
# no openshell.ai/sandbox-runtime-bootstrapping annotation
# (no sandbox pod in the namespace)
$ openshell sandbox start <name>
Error: sandbox must be Stopped, Completed, or a failed main-process Error to start (current phase: Provisioning)
$ openshell sandbox stop <name>
Error: sandbox must be Ready to stop (current phase: Provisioning)
Root cause (pointers; as of e1f3c82ca)
- The record's
status.phaseisProvisioningbut the backend CR is empty.
start_sandboxrejects it because the gate only admits
Stopped | Completed | Starting | is_failed_main_process_result | provisioning_timeout
(crates/openshell-server/src/compute/mod.rs,start_sandbox, ~L1784). - The reconcile sweep re-derives phase from the driver snapshot; with no pod /
replicas: 0, theProvisioningcase falls through to the_ => phasedefault
(~L6564, fed byderive_phase~L6863), so the stale value is kept. suspend_sandbox_runtime_after_dependency_failure
(crates/openshell-driver-kubernetes/src/driver.rs, ~L4111) is what drives
replicas: 0on the dependency-failure path.
Expected
One of:
- The reconcile loop corrects a
Provisioningrecord whose backend reports the
sandbox as empty (no pod /replicas: 0) instead of perpetuating it — cf. #2933, which
madederive_phaseletReadywin over a staleSuspended; or startaccepts a staleProvisioningwhose backend is empty and routes it through the
existing idempotent-restart path — reusing the persisted identity and re-attaching the
existing workspace PVC — the same way theStartingbranch already does (~L1800).
Actual
No recovery short of delete + recreate. An out-of-band edit of the gateway store
(flipping Sandbox.status.phase from Provisioning to Stopped) does make start
succeed and preserves the workspace — confirming this is a fixable record desync rather
than data loss — but it is not a supported operator path.
Proposed direction
- Extend #4231's recovery approach to the
Provisioning-with-empty-backend case (not
justError), or - correct the stale
Provisioninginderive_phase/reconcile (precedent: #2933).
Either way, start from this state should reuse the persisted identity and the existing
workspace PVC (idempotent restart), not require a delete.
References
- #3308 — unrecoverable
Error, stop/start refuse (same symptom class) - #4209 — gateway restart → unrecoverable
Error(same trigger family) - #4231 — open: allow stopping errored sandboxes for recovery (in-flight recovery mechanism)
- #2932 / #2933 — stale condition perpetuated by reconcile; fix in
derive_phase(closest mechanism) - #4078 — serialize lifecycle cleanup with sandbox restart (related stop/start race)
- Lingua principale
- Rust
- Stelle
- 15.4k
- Fork
- 1.7k
- Merge medio
- 1g 21h
- PR unite (30g)
- 366
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di NVIDIA/OpenShell
-
state:triage-needed
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
I maintainer di solito rispondono entro 1 giorno
-
docs: document workspace and provider label capabilitiesForse già presa @johntmyers l’ha presa 3 giorni fa. Apertaarea:docs
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
NVIDIA/OpenShell#4250 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
bug(driver-mxc): test helper fails to compile after gateway-name argumentForse già presa @feloy l’ha presa 5 giorni fa. Apertastate:triage-needed
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
bug: install.sh ignores XDG_CONFIG_HOME for the local gateway configForse già presa @fede-kamel l’ha presa 8 giorni fa. Apertaarea:cli os:linux os:macos state:validated
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
NVIDIA/OpenShell#4042 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
state:triage-needed
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
NVIDIA/OpenShell#3995 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di NVIDIA/OpenShell
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
rescript-lang/rescript#8765 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
farion1231/cc-switch#8072 ·
I maintainer di solito rispondono entro 1 giorno
-
Python 3.15 supportForse già presa @amnesiaof l’ha presa oggi. ApertaL: python L: python:uv
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
dependabot/dependabot-core#16524 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
I maintainer di solito rispondono entro 1 giorno
-
[Bug]: Migration link in chromadb/config.py error message returns 404Forse già presa @Imad2702 l’ha presa oggi. Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
chroma-core/chroma#7879 ·
I maintainer di solito rispondono entro 1 giorno