gateway: robustness hardening plan for WaveKV-synced state (fault isolation, blast-radius control)
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Facilidade para iniciantes
- 25/100
- Tipo de issue
- Funcionalidade
- Clareza
- Precisa de esclarecimento
- Status de atividade
- Pouca atividade
- Stack de tecnologia
- rust
- Domínio
- backend, databases, distributed-systems, observability, security
Direção de pesquisa
Comece com a sequência P0 sugerida, lendo gateway/src/web_routes/wavekv_sync.rs, gateway/src/kv/, gateway/src/main_service.rs e a crate wavekv; use gateway/test-run/e2e/ para a matriz de versões mistas. A tarefa estará concluída quando as salvaguardas listadas de validação, contenção, recuperação, observabilidade e sincronização estiverem implementadas, as verificações relevantes passarem e nenhum registro inválido derrubar estado não relacionado.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Context
dstack-gateway replicates its state (instances, nodes, certificates, DNS credentials, ACME account) across nodes via WaveKV. The gateway runs in a TEE and must stay robust when the KV contains bad data: a corrupt or malformed record for one CVM must never take down other instances or the whole cluster.
This issue tracks the findings of a robustness review of gateway/src/kv/, gateway/src/web_routes/wavekv_sync.rs, the KV→ProxyState consumption paths in gateway/src/main_service.rs, and the wavekv crate itself. The wavekv protocol redesign that came out of the same review is tracked separately in Phala-Network/wavekv#2 (RFC 0001: delta-state synchronization); the migration-relevant parts are summarized at the bottom.
What is already right (keep): ProxyState is the primary read path so the data plane does not depend on KV availability; per-key decode failures are warn-and-skip; get_acme_credentials() fails closed on corruption; peer URLs are validated; the sync endpoint enforces RA-TLS same-app-id mTLS, a 16 MiB body cap, and rejects node_id == 0.
P0 — paths where one bad input becomes a global outage
1. Decompression bomb on the sync endpoint
web_routes/wavekv_sync.rs: data.open(16.mebibytes()) caps the compressed size only; GzDecoder::read_to_end is unbounded. Gzip expands up to ~1000:1, so a single 16 MiB request from a compromised or buggy peer can force a ~16 GiB allocation → OOM kill.
- Cap decompressed size (
decoder.take(limit)+ error on overflow); also cap entry count and per-entry value size in the decodedSyncMessage.
2. One bad instance record can break the whole WireGuard config
reload_instances_from_kv_store imports KV instances into ProxyState verbatim; reconfigure() renders templates/wg.conf with escape = "none" and runs wg syncconf.
- A malformed
public_keymakeswg syncconfreject the entire config file — all instances lose wg updates (error is only logged). A key containing newlines can injectEndpoint=/AllowedIPs=directives. valid_ip()(client range / broadcast / reserved nets) is enforced at registration, but not on the KV import path: a synced instance can carry the gateway's own wg IP, a reserved-net IP, or an IP duplicating another instance (LWW cannot enforce cross-key invariants).- Public-key uniqueness is checked in
new_client_by_idbut not on import. - Re-run all registration-path semantic validation at the KV→ProxyState import boundary (pubkey is valid base64 32 bytes; IP in range; IP/pubkey uniqueness); skip only the offending instance, never abort the batch.
- Assert key/IP formats once more before rendering
wg.conf.
3. unwrap_or_default() on corrupt global keys silently changes global behavior
The fail-closed pattern of get_acme_credentials() (absent ≠ tombstone ≠ corrupt) is not applied to its siblings:
get_certbot_config(): a corrupt record silently falls back to defaults, i.e. switchesacme_urlto Let's Encrypt production and resets renewal intervals.get_default_dns_credential_id,get_acme_attestation, etc. treat corruption as absence.- Extract the three-state helper (missing / deleted / corrupt→
Err) and apply it to every global key whose corruption must not silently change behavior; alert on the corrupt case.
4. Wall-clock LWW + clock skew: one bad write can poison state cluster-wide
- A node with a future clock (or a corrupted
i64::MAX-ish timestamp) wins every LWW conflict; the key becomes unfixable until real time catches up. No admin override exists. - Gateway aggregations take
maxacross nodes:get_instance_latest_handshake/get_node_latest_last_seen. One node writing future handshake timestamps keeps dead CVMs "alive" cluster-wide —recycle()never fires and top-N routing is distorted. - Clamp/reject timestamps beyond
local_now + max_drifton ingest and in aggregations. - Admin "force put" escape hatch (writes with
max(existing.ts)+1). - (Long-term: HLC in wavekv v2 — see RFC.)
5. Local WAL/snapshot corruption prevents startup, though state is fully replicated
Node::new_with_persistence: read_all_ops() hard-fails on a checksum/deserialize error (while find_last_sequence tolerates the same); a corrupt 4-byte length prefix can trigger a multi-GiB allocation; a corrupt snapshot is fatal with no .bak fallback. Torn WAL tails are the normal crash artifact and should never brick the gateway.
- Wrap
KvStore::new: on init failure, quarantine the data dir (rename.corrupt), start empty, re-bootstrap from peers. - Upstream (wavekv): tolerant tail recovery, length sanity checks, snapshot generation fallback — tracked in RFC §3.10.
6. Global-lock amplifiers
Node::read/writeuses.expect()on a poisonablestd::sync::RwLock: any panic while holding the write lock turns into a permanent crash loop. Switch toparking_lot(upstream).persist_if_dirtyserializes + fsyncs the snapshot inside the global write lock, stalling registration/sync paths as state grows. CloneCoreStateunder the lock, write outside.
P1 — containment and correctness
7. "Decode failure = silently invisible" needs quarantine + visibility
A corrupt inst/ record makes that CVM vanish from routing with only a warn log.
- Maintain a quarantine list (key → error → entry meta) surfaced via admin/status; per-prefix decode-failure metrics + alerting.
8. No schema/quota enforcement on ingest
Any same-app peer can replicate arbitrary keys of arbitrary size to every node, persisted forever.
- Key-prefix whitelist (
inst/,node/,cert/,dns_cred/,global/,__peer_addr/,conn/,handshake/,last_seen/), per-prefix value-size caps, global key-count/byte caps. - GC for
cert/{domain}/attestation/{timestamp}history (unbounded today). - Note:
cleanup_expired_tombstonesis never called — tombstones grow forever; if ever enabled, it must be watermark-coordinated first (resurrection risk; see RFC §6).
9. Remote deletions never remove instances from ProxyState
reload_instances_from_kv_store only upserts. An instance recycled on node A stays in node B's ProxyState/wg config until B's own recycle timeout — a deregistered CVM remains routable in the window.
- On reload, explicitly remove instances present locally but absent/tombstoned in KV (with a reg_time grace window for not-yet-synced local registrations).
10. Schema evolution in a mixed-version cluster
Values use rmp_serde positional encoding. New-data→old-decoder fails during rolling upgrades/rollbacks → instances silently vanish on old nodes (compounds item 7).
- Written policy: additive
#[serde(default)]fields only; bidirectional decode compat tests with old-version fixtures. - Consider
to_vec_namedor a{version, body}envelope for new key types.
11. Minor consistency checks
-
list_zt_domain_configs: assert value.domain matches the domain in the key; mismatch → quarantine. - Document the worst-case behavior of the best-effort LWW locks (
cert/{domain}/lock, rotation lock) — concurrent renewal is bounded by ACME idempotency today. - Per-peer sync-lag metrics (a peer stuck in gap-drop today only produces warn spam).
Design-level follow-ups
- Formalize the import boundary: a single module through which all KV→ProxyState/CertStore/wg data flows, with decode → semantic → invariant validation layers and quarantine on failure (gives items 2/7/9/11 one home).
- Failure-domain separation: routing state, cert/keys, DNS creds, and attestation history share one lock/WAL/snapshot; consider per-domain stores so a poisoned cert subtree cannot stall instance sync.
- Secrets in KV: cert private keys, Cloudflare tokens, and the ACME account key are replicated in plaintext to every peer and included in full-dump responses. Consistent with the RA-TLS same-app trust model, but it makes any single gateway compromise a total credential compromise. Audit debug/admin endpoints for raw-KV exposure; consider KMS-derived encryption for sensitive values.
- Observability minimum set: per-prefix decode failures, quarantined keys, per-peer sync lag/last-success, KV size,
wg syncconffailure alert.
WaveKV upgrade plan (tracked in Phala-Network/wavekv#2)
The protocol-level items (delta-state sync replacing op-logs; state digest for silent-divergence detection; ingest admission hooks/quotas; WAL/snapshot recovery hardening; coordinated tombstone GC; HLC deferred) are specified in the RFC. Gateway-side integration follows its staged plan:
- Phase 0 (wavekv 1.x prep): adopt
state_digest()+ per-peer sync metrics; compare digests via the admin plane; gate = digest equality across the production cluster. - Phase 1 (wavekv 2.0 dual-stack): rolling upgrade one CVM at a time; gateway adds the
/wavekv/sync2/{store}route with probe-and-fallback negotiation; promotion gate per node = cluster-wide digest equality + shim counters clean. Any single node can roll back (snapshot/WAL formats stay v1-loadable). - Phase 2 (cleanup): after ≥14 days of digest equality on all-v2, drop the v1 sync route; only then take wire-breaking follow-ups (HLC).
- Mixed-version e2e matrix runs on
gateway/test-run/e2e/(3-node harness); test list in RFC §8.5.
Suggested sequencing overall: P0.1 → P0.2 → P0.5 → P0.3+P0.4, then the import-boundary refactor + observability, with the wavekv Phase 0/1 work proceeding in parallel under the RFC.
- Linguagem predominante
- Rust
- Estrelas
- 551
- Forks
- 97
- Merge médio
- 1d 8h
- PRs com merge (30d)
- 182
Preparar o ambiente
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de Dstack-TEE/dstack
-
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 35/100
Dstack-TEE/dstack#1384 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 30/100
Dstack-TEE/dstack#1301 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 3/5 1-2 dias Facilidade para iniciantes 55/100
Dstack-TEE/dstack#1300 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 48/100
Dstack-TEE/dstack#1299 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 48/100
Dstack-TEE/dstack#1298 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de Dstack-TEE/dstack
Issues semelhantes
-
`categorize_command` has no `uv` arm, so every `rtk uv …` row counts as `other` in the ecosystem mixAbertaarea:api bug good first issue priority:low
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 92/100
rtk-ai/rtk#4316 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 92/100
Mantenedores costumam responder em até 1 dia
-
area/cli kind/bug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 90/100
Mantenedores costumam responder em até 1 dia
-
enhancement
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
-
good first issue open-endedness: low type: new feature
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100