[Feature]: Support PostgreSQL-backed shared spend state for zero-downtime rolling updates
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- kubernetes, postgresql, typescript
- Domain
- backend, databases, devops, distributed-systems
Research direction
Start by reading src/lib/spend-reservation-ledger.ts and src/lib/spend-ledger-owner.ts, then review the existing /readyz behavior and issues #5123 and #4173. Define the supported shared-state topology, migration and rollback checks, and failure behavior before implementation. Done means verified two-pod Ready-before-stop operation with shared admission limits and safe fail-closed behavior.
Written by the indexing model from the issue text.
Description
Area
Service lifecycle
What are you trying to accomplish?
Run OpenCodex as a service that can perform a Kubernetes Ready-before-stop rolling replacement (maxSurge: 1, maxUnavailable: 0). During the overlap, the old pod must continue serving requests until the replacement is actually ready, while both instances obey one logical spend/reservation state and the same admission limits. This is an opt-in multi-instance deployment mode; ordinary single-instance installations should not need PostgreSQL.
What prevents this today?
The current supported topology is one live writer per OPENCODEX_HOME. The durable spend ledger is a JSONL journal, while a separate SQLite spend-ledger-owner.sqlite holds a process-lifetime BEGIN IMMEDIATE ownership transaction. The replacement pod using the same home cannot acquire that lease while the old pod runs, so it cannot become Ready before the old pod stops. This is correct fail-closed behavior for the current topology, not a broken SQLite lock. Recreate avoids concurrent writers but interrupts service; separate homes allow overlap but do not by themselves provide one shared transactional reservation/budget state.
The source explicitly notes that multi-host shared storage requires a distributed transaction boundary and is outside the local lease's scope: spend-reservation-ledger.ts, topology contract; current owner implementation.
What should OpenCodex do?
Please consider an opt-in PostgreSQL-backed shared-state mode for multi-pod operation, rather than replacing the default local JSONL ledger and SQLite ownership lease for everyone. A supported mode should:
- Persist reservations, settlements and recovery state through a transactional shared backend, with cross-pod coordination/fencing appropriate to concurrent admission decisions; do not merely move the SQLite lease to PostgreSQL while still serializing ownership across both pods for their entire lifetimes.
- Let a replacement become Ready while the old instance is serving, without either pod making conflicting reservations or exceeding a shared ceiling. Readiness should reflect the ability to safely serve under the selected backend; a failed connection/coordination attempt must fail closed and report a bounded, non-sensitive diagnostic.
- Define the behavior for in-flight requests and graceful drain, crashes, restart/replay, database unavailability, and a pod losing its right to admit requests.
- Document the scope of other mutable
OPENCODEX_HOMEstate that must be shared, isolated, or migrated before claiming general multi-pod support; a PostgreSQL spend ledger alone may not make the full runtime safe to overlap. - Keep the existing local single-writer deployment as the default, and document initial migration, rollback/downgrade and operational checks for the opt-in mode.
Could maintainers advise whether PostgreSQL is the right distributed transaction boundary, and what supported multi-instance state topology would be required? The desired outcome is verified Ready-before-stop behavior, not a database swap for its own sake.
Example usage or interface
Illustrative proposed interface, not an existing OpenCodex option (exact names and implementation are for maintainers to decide):
# OpenCodex configuration, illustrative only
spendLedger:
backend: postgres
connectionEnv: OPENCODEX_SPEND_LEDGER_DATABASE_URL
# Kubernetes Deployment
strategy:
type: RollingUpdate
rollingUpdate:
maxSurge: 1
maxUnavailable: 0
Expected observation: with two pods temporarily running, the replacement reaches /readyz before the old pod is removed, and concurrent requests across both pods respect one shared limit. If shared-state safety cannot be established, the replacement remains unready and the old pod continues serving.
Alternatives or workarounds
- Shared home with the current local lease: correctly rejects the second writer, so Ready-before-stop cannot complete.
- Separate homes: avoids the local lease conflict but does not establish shared reservations or solve other state-consistency questions.
Recreate/manual stop-then-start: preserves single-writer semantics at the cost of downtime.- A graceful drain/update command alone: valuable for in-flight requests but cannot make two instances concurrently safe against the same spend state.
Additional context
#5123 established the local single-writer topology and explicitly deferred a distributed transaction boundary; this proposal asks for an additional supported topology, not a reversal of that decision. #4173 concerns upgrade/drain coordination but does not provide a multi-instance shared-state backend. This is a source-based deployment limitation, not a claim that a two-pod live experiment or PostgreSQL implementation has already been validated.
Checks
- I searched existing issues and documentation.
- This request describes a concrete OpenCodex workflow rather than merely naming a desired technology.
- I removed secrets and personal data.
- Dominant language
- TypeScript
- Stars
- 16.9k
- Forks
- 1.3k
- Avg merge
- 4h 17m
- Merged PRs (30d)
- 594
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from lidge-jun/opencodex
-
provider provider-compatibility
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
bug proxy tools
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
lidge-jun/opencodex#6648 · 1 comment ·
Maintainers usually reply within 1 day
-
bug proxy
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
lidge-jun/opencodex#6646 · 2 comments ·
Maintainers usually reply within 1 day
-
bug service
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
lidge-jun/opencodex#6643 · 1 comment ·
Maintainers usually reply within 1 day
-
bug proxy service
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
lidge-jun/opencodex#6642 · 1 comment ·
Maintainers usually reply within 1 day
All issues in lidge-jun/opencodex
Similar issues
-
bot:ai-assisted status:untriaged
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
midnightntwrk/midnight-js#1424 ·
Maintainers usually reply within 1 day
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
mksglu/context-mode#1268 ·
Maintainers usually reply within 5 days
-
[bug] Setup fails with "Cannot find matching keyid" when an older Node's corepack is on PATHPossibly taken @EyalPoly claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
MystenLabs/MemWal#1124 · 2 comments ·
Maintainers usually reply within 1 day
-
Edit:Opencheck:failed streams:edit
Difficulty 2/5 1-3 hours Newbie friendliness 60/100
iptv-org/iptv#54352 · 1 comment ·
Maintainers usually reply within 1 day