[Feature]: complete the OAuth account-pool lifecycle — session affinity, 401/403 rotation, pool health, and stable reset-credit identity
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- typescript
- Domain
- api, authentication, backend, frontend
Research direction
This is a multi-track account-pool feature spanning src/oauth/generic-account-failover.ts, src/providers/quota.ts, gui/src/components/provider-workspace/ProviderAuthPanel.tsx, src/codex/reset-credit-operation-ledger.ts, src/codex/auth-api.ts, and src/codex/warmup.ts. Read the existing failover and ledger implementations first, then separate the session, rotation, health, identity, and warmup requirements. Done means the configured lifecycle, aggregate attribution, stable reset operation IDs, and durable one-shot warmup all work together.
Written by the indexing model from the issue text.
Description
CLI truthfulness slice landed; pool epic remains open
#3797 (a92835dbc6) makes inert generic-pool thresholds explicit in CLI status. This does not complete the broader account-pool lifecycle epic. Existing reset-operation identity work must not be counted as unimplemented solely from the older summary.
Verified against dev 5759d9ea2f1e7281cdc01eb9628f2e0a123fb59c. The attribution record was added in #3811.
Area
Authentication and account pool
What are you trying to accomplish?
Run OpenCodex against a pool of OAuth/subscription accounts and have it pick the right account, keep a conversation on that account, and recover from per-account failures without operator intervention.
This consolidates the surviving scope from four reports whose first tranche already landed. Each is closed individually and absorbed here so the remaining work stays visible in one place:
- #695 (@luwei1990) — generic OAuth account-pool failover with session affinity and quota-aware selection
- #1062 (@agentHits) — Google Antigravity multi-account management, pool health and quota UX
- #1977 (@dbc-hbin) — pinning the re-initialization window of zero-usage Codex accounts
- #2275 (@luvs01) — durable operation identity for manual reset-credit retries
What prevents this today?
Partial implementations exist and work, but four concrete gaps remain.
Session affinity is not reusable. src/oauth/generic-account-failover.ts implements 429 cooldown, rotation, and quota-ranked initial selection (landed in 816f3a159). It does not keep a multi-turn conversation pinned to the account that started it, so a long session can drift across accounts mid-conversation.
Failover classes are too narrow. Only 429 rotates. A 401 and a provider-classified 403 are terminal for the pool even when another account in the pool would succeed.
Pool-level health and attribution are missing. Per-account quota is surfaced (src/providers/quota.ts, gui/src/components/provider-workspace/ProviderAuthPanel.tsx), but there is no aggregate pool capacity view and no per-account attribution of usage, so an operator cannot see which account is carrying the load or why a selection was made.
Manual reset-credit retries still mint a new identity per request. The durable ledger landed (src/codex/reset-credit-operation-ledger.ts, 7c68768ca), but the manual consume endpoint generates a fresh UUID per call at src/codex/auth-api.ts, so a retried manual reset is not recognized as the same logical operation the ledger was built to track.
Zero-usage accounts have no durable warmup scheduling. src/codex/warmup.ts provides a manual warmup primitive only; there is no one-shot scheduled shadow request to pin an account's re-initialization window.
What should OpenCodex do?
- Keep a conversation on its selected account for the life of the session, with an explicit, configurable release condition, and fall over only when that account genuinely cannot serve.
- Treat 401 and provider-classified 403 as rotatable within the pool, with the credential marked unhealthy rather than the request failed.
- Expose aggregate pool health — capacity, per-account quota state, and which account served a request — in the dashboard and in request history.
- Accept a caller-supplied stable operation ID on the manual reset-credit endpoint and CLI, and open/settle that exact ID in the existing ledger so a retry is idempotent.
- Allow a durable one-shot warmup to be scheduled for a zero-usage account so its window starts predictably.
Example usage or interface
// config.json
{
"providers": {
"example": {
"accountPool": {
"strategy": "quota",
"sessionAffinity": "conversation", // pin a turn chain to one account
"rotateOn": ["429", "401", "403:quota"] // classes that rotate instead of failing
}
}
}
}
# retrying this must settle the SAME ledger operation, not open a second one
ocx codex reset-credit --account [email protected] --operation-id 6f1c1f2e-0f2a-4a1e-9a1b-0d2f7c8e5b31
Alternatives or workarounds
Operators currently work around this by running one account per proxy instance and load-balancing externally, which defeats quota-aware selection and makes pool health invisible.
Additional context
Landed prior work referenced above: 816f3a159 (generic account failover), 7c68768ca (reset-credit operation ledger).
Absorbs #695, #1062, #1977, #2275. Credit for the original analysis belongs to @luwei1990, @agentHits, @dbc-hbin, and @luvs01.
Checks
- I searched existing issues and documentation.
- This request describes a concrete OpenCodex workflow rather than merely naming a desired technology.
- I removed secrets and personal data.
- Dominant language
- TypeScript
- Stars
- 16.9k
- Forks
- 1.3k
- Avg merge
- 4h 15m
- Merged PRs (30d)
- 607
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from lidge-jun/opencodex
-
bug proxy tools
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
lidge-jun/opencodex#6648 · 1 comment ·
Maintainers usually reply within 1 day
-
bug proxy
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
lidge-jun/opencodex#6646 · 2 comments ·
Maintainers usually reply within 1 day
-
[Bug] The restart drain fence refuses new requests for the full 60s while Codex exhausts its retries in about six secondsPossibly taken A pull request linked to this issue is open or already merged. Openbug service
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
lidge-jun/opencodex#6643 · 1 comment ·
Maintainers usually reply within 1 day
-
account-pool bug catalog proxy
Difficulty 4/5 3-5 days Newbie friendliness 52/100
Maintainers usually reply within 1 day
-
account-pool bug gui
Difficulty 3/5 Half a day Newbie friendliness 78/100
lidge-jun/opencodex#6701 · 1 comment ·
Maintainers usually reply within 1 day
All issues in lidge-jun/opencodex
Similar issues
-
Bump Firebase JS SDK (12.19.0 → 13.0.0)Possibly taken @SelaseKay claimed this today. OpenNeeds Attention type: enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
invertase/react-native-firebase#9364 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 4 days
-
e2e-failure ready-to-code
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
redhat-developer/rhdh-plugin-export-overlays#4261 · 1 comment ·
Maintainers usually reply within 1 day
-
[Bug] 官网文档的图片挂了Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
area:cli bug triage:in-progress
Difficulty 1/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day