Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

bug(kubernetes): sandbox stuck at `Provisioning` with `replicas: 0` has no recovery path (`start`/`stop` both refuse)

Aperta
#4,371 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
30/100
Tipo di issue
Bug
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
kubernetes, rust
Ambito
backend

Direzione di ricerca

Start with the phase gate in start_sandbox in crates/openshell-server/src/compute/mod.rs (~L1784) and the reconcile fallthrough near derive_phase (~L6863), then see how crates/openshell-driver-kubernetes/src/driver.rs suspend_sandbox_runtime_after_dependency_failure (~L4111) sets replicas to 0. Done means a sandbox stuck in Provisioning with an empty backend either gets corrected by reconcile or can be started through the idempotent-restart path, with the workspace PVC kept. The issue offers two directions, so agree the approach with maintainers before coding.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

state:triage-needed

Summary

A sandbox can end up with the gateway's Sandbox record at phase Provisioning
while the backend Sandbox CR is at spec.replicas: 0 with no pod. In that state
there is no supported recovery path:

  • sandbox start is rejected — the phase gate only admits Stopped,
    Completed, Starting, a failed-main-process Error, or a provisioning timeout.
  • sandbox stop is rejected — it requires Ready.
  • The reconcile sweep does not correct the record: with replicas: 0 / no pod, the
    Provisioning case falls through to the reconcile/derive_phase default and the
    stale value is perpetuated rather than corrected.

The only exits are delete + recreate, or out-of-band edits to the gateway's store.

This is a distinct variant of the "unrecoverable phase" bugs #3308
and #4209, which are Error-phase. Here the record is
Provisioning, so the recovery work in #4231 (allow stop from
Error) does not reach it, and start still rejects it. The closest mechanism match is
#2932 (a stale condition perpetuated by reconcile), fixed by
#2933.

Environment

  • OpenShell: 0.1.2 (gateway ghcr.io/nvidia/openshell/gateway:0.1.2, supervisor 0.1.2, CLI 0.1.2)
  • Driver: Kubernetes (agent-sandbox v0.4.6, CRD sandboxes.agents.x-k8s.io v1alpha1)
  • Kubernetes: v1.27.3 (kind)

Reproduction

Non-deterministic (a race). Observed triggers:

  1. A node/machine reboot while the gateway is bootstrapping a sandbox — the gateway pod
    restarts during cluster stabilization, and a transient dependency failure fires
    suspend_sandbox_runtime_after_dependency_failure, resetting replicas: 0 while the
    record is still Provisioning.
  2. A stop immediately followed by start, leaving the record mid-transition.

Observed state

$ openshell sandbox list
NAME   ...   Provisioning

$ kubectl get sandboxes -n <ns> -o yaml
# spec.replicas: 0
# no openshell.ai/sandbox-runtime-bootstrapping annotation
# (no sandbox pod in the namespace)

$ openshell sandbox start <name>
Error: sandbox must be Stopped, Completed, or a failed main-process Error to start (current phase: Provisioning)

$ openshell sandbox stop <name>
Error: sandbox must be Ready to stop (current phase: Provisioning)

Root cause (pointers; as of e1f3c82ca)

  • The record's status.phase is Provisioning but the backend CR is empty.
    start_sandbox rejects it because the gate only admits
    Stopped | Completed | Starting | is_failed_main_process_result | provisioning_timeout
    (crates/openshell-server/src/compute/mod.rs, start_sandbox, ~L1784).
  • The reconcile sweep re-derives phase from the driver snapshot; with no pod /
    replicas: 0, the Provisioning case falls through to the _ => phase default
    (~L6564, fed by derive_phase ~L6863), so the stale value is kept.
  • suspend_sandbox_runtime_after_dependency_failure
    (crates/openshell-driver-kubernetes/src/driver.rs, ~L4111) is what drives
    replicas: 0 on the dependency-failure path.

Expected

One of:

  1. The reconcile loop corrects a Provisioning record whose backend reports the
    sandbox as empty (no pod / replicas: 0) instead of perpetuating it — cf. #2933, which
    made derive_phase let Ready win over a stale Suspended; or
  2. start accepts a stale Provisioning whose backend is empty and routes it through the
    existing idempotent-restart path — reusing the persisted identity and re-attaching the
    existing workspace PVC — the same way the Starting branch already does (~L1800).

Actual

No recovery short of delete + recreate. An out-of-band edit of the gateway store
(flipping Sandbox.status.phase from Provisioning to Stopped) does make start
succeed and preserves the workspace — confirming this is a fixable record desync rather
than data loss — but it is not a supported operator path.

Proposed direction

  • Extend #4231's recovery approach to the Provisioning-with-empty-backend case (not
    just Error), or
  • correct the stale Provisioning in derive_phase/reconcile (precedent: #2933).

Either way, start from this state should reuse the persisted identity and the existing
workspace PVC (idempotent restart), not require a delete.

References

  • #3308 — unrecoverable Error, stop/start refuse (same symptom class)
  • #4209 — gateway restart → unrecoverable Error (same trigger family)
  • #4231 — open: allow stopping errored sandboxes for recovery (in-flight recovery mechanism)
  • #2932 / #2933 — stale condition perpetuated by reconcile; fix in derive_phase (closest mechanism)
  • #4078 — serialize lifecycle cleanup with sandbox restart (related stop/start race)
Lingua principale
Rust
Stelle
15.4k
Fork
1.7k
Merge medio
1g 21h
PR unite (30g)
366

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di NVIDIA/OpenShell

Tutte le issue di NVIDIA/OpenShell

Issue simili

Altre issue su Rust

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.