feat(ha): add production scaling signals and graceful gateway redistribution
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- helm, kubernetes, postgresql, rust
Direzione di ricerca
Start with PR #1868 and its linked discussions on graceful redistribution, scaling metrics, advisory-lock serialization, and WatchSandbox polling. Trace the gateway lifecycle, supported metrics endpoint, Helm documentation, mutation locking, and cross-replica watch synchronization. Done means the acceptance criteria and load or integration tests cover all four production-scale behaviors.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
User Story
As an OpenShell cluster operator, I want multi-replica gateways to expose actionable capacity signals and redistribute supervisor ownership during planned changes, so that scaling and rolling upgrades do not leave one replica overloaded while new replicas remain idle.
Problem Statement
PR #1868 makes multi-replica gateway routing possible, but several production-scale behaviors remain incomplete:
- Supervisor ownership stays with the original connection. Rolling restarts and failover can leave the last surviving replica holding most sessions until sandboxes naturally reconnect.
- The gateway does not expose held-session, pending-relay, or peer-relay rate and latency metrics suitable for capacity planning or autoscaling.
- Cross-object mutations use one fleet-wide PostgreSQL advisory-lock key, so unrelated sandbox and provider mutations serialize across all replicas.
WatchSandboxperforms one database lookup per actively watched sandbox, per polling interval, per replica.
Impact / Why This Matters
Operators can add gateway replicas, but cannot tell when capacity is exhausted or rely on replicas receiving a balanced share of established supervisor sessions. The current workaround is to overprovision, manually inspect logs, and wait for sandbox churn or force reconnects after a rollout. That does not prevent pending-relay exhaustion, database amplification, or fleet-wide mutation contention at larger scale.
Proposed Design
Provide an observable, documented HA lifecycle with graceful draining and bounded shared-database load:
- A terminating gateway stops accepting new ownership, releases or transfers existing ownership, and prompts affected supervisors to reconnect before the termination grace period expires.
- Gateway metrics expose held supervisor sessions, pending relays versus capacity, and peer-relay request rate, failure rate, and latency.
- The Helm chart documents those signals and provides an optional autoscaling configuration, or a documented integration contract for a Kubernetes metrics adapter.
- Unrelated object mutations do not share one fleet-wide serialization point when their invariants cannot overlap.
- Cross-replica sandbox watches use a batched or notification-based mechanism whose database work does not grow as one query per watched sandbox on every replica.
Acceptance Criteria
- During a graceful gateway rollout, established supervisor ownership redistributes without waiting for natural sandbox churn, within the configured termination grace period.
- Operators can observe per-replica held sessions, pending relays/capacity, and peer-relay rate, failures, and latency through the supported metrics endpoint.
- The Helm documentation explains how to use these signals for horizontal scaling; any chart-provided HPA behavior is optional and configurable.
- Mutations for unrelated sandboxes or providers can proceed concurrently across replicas unless they share a documented cross-object invariant.
- Cross-replica watch synchronization avoids one database query per watched sandbox per interval per replica.
- Load or integration tests cover rollout redistribution, relay saturation signals, concurrent unrelated mutations, and large watch sets.
- HA architecture and Kubernetes operator documentation describe the lifecycle, capacity signals, and database-scaling behavior.
Alternatives Considered
Keeping ownership sticky is simple and preserves correctness, but leaves rollout skew entirely dependent on sandbox churn. Relying only on logs does not provide stable autoscaling inputs. Increasing the PostgreSQL connection pool masks neither fleet-wide lock serialization nor query amplification.
Agent Investigation
This groups the remaining production-scale review discussions from PR #1868:
- Graceful redistribution and sticky ownership: https://github.com/NVIDIA/OpenShell/pull/1868#discussion_r3945390329
- Scaling metrics and HPA signals: https://github.com/NVIDIA/OpenShell/pull/1868#discussion_r3945372631
- Fleet-wide advisory-lock serialization: https://github.com/NVIDIA/OpenShell/pull/1868#discussion_r3946935300
- Watch-poller query amplification: https://github.com/NVIDIA/OpenShell/pull/1868#discussion_r3952505962
Related umbrella issue: #1021.
- Lingua principale
- Rust
- Stelle
- 8.7k
- Fork
- 1.3k
- Merge medio
- 2g 8h
- PR unite (30g)
- 271
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di NVIDIA/OpenShell
-
area:docs
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
-
state:triage-needed
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
-
area:cli state:validated
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
-
state:triage-needed
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
-
area:build spike state:review-ready state:stale
Difficoltà 2/5 Mezza giornata Idoneità per principianti 68/100
Tutte le issue di NVIDIA/OpenShell
Issue simili
-
Browser (wasm) relay client cannot connect to relays whose URL has a trailing-dot FQDN hostname Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
n0-computer/iroh#4550 ·
-
impl detach for native Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
paritytech/zombienet-sdk#591 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
farion1231/cc-switch#7638 · 1 commento ·
-
onnx-ir re-exports ModelProto and GraphProto but not NodeProto, AttributeProto and AttributeType Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100