[Bug] The restart drain fence refuses new requests for the full 60s while Codex exhausts its retries in about six seconds
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 70/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- typescript
- Domain
- backend
Research direction
Start in src/server/management/system-restart.ts at the deadline calculation around line 366 and the drain call around line 393. Check how the restart handoff uses the remaining deadline, then verify that bounding the refusal window still lets the drain and replacement startup proceed. Done means new requests are refused for no longer than the handoff cap and a client can retry against the replacement.
Written by the indexing model from the issue text.
Description
Client or integration
Codex CLI
Area
Service lifecycle
Summary
The restart handoff drains for up to MEMORY_DRAIN_RESTART_MS (60s). The intent is to let an in-flight turn finish, but during that whole window every NEW data-plane request is refused, and Codex spends its five sampling retries in roughly six seconds. So the client gives up well before the replacement is listening, and the turn fails even though the restart itself is quick.
Measured on this machine (2026-09-26): with the 60s default, a 42-second fence failed the turn with unexpected status 503 - and that was already after the refusal stopped being a fatal capacity error (see the related issue below). A two-second cap, plus the replacement's roughly one-second bind gap, passed. In other words the fence length, not the refusal itself, is what the client runs out of.
Reproduction
ocx start --port 10100.- Start a Codex turn.
- Trigger the drain-and-restart path (dashboard
POST /api/system/restart, orocx service restart) while the turn is in flight. - Watch the turn: the client retries against a listener that is refusing for the full fence, exhausts its retries, and ends with
unexpected status 503instead of continuing against the replacement.
Version
2.78.0
Operating system
macOS 27.0.1 (arm64). The fence is shared code, so this should apply on every platform.
Provider and model
Not provider-specific.
Logs or error output
Source location: src/server/management/system-restart.ts - the deadline is set at line 366 (const restartDeadlineMs = now() + MEMORY_DRAIN_RESTART_MS;) and consumed at line 393 (const remainingMs = Math.max(0, restartDeadlineMs - now());), which is then handed to drain(undefined, remainingMs).
Suggested change
Bound the restart fence rather than the drain - keep the drain free to finish an in-flight turn, but stop refusing new requests for longer than the client can retry:
const RESTART_HANDOFF_INFLIGHT_WAIT_MS = 2_000;
const remainingMs = Math.min(
Math.max(0, restartDeadlineMs - now()),
RESTART_HANDOFF_INFLIGHT_WAIT_MS,
);
A turn cut by this cap is retried by the client against the replacement that the same handoff starts, which is why two seconds is enough: the gap it has to cover is the replacement's bind, not the full drain.
This is what we run as a local patch. Happy to open a PR if you would rather review it that way.
Related: the code carried by the refusal is the other half of this handoff problem, filed separately as #6642.
Screenshots and supporting files
None.
Redacted configuration
{ "hostname": "127.0.0.1", "port": 10100 }
Checks
- I searched existing issues and documentation.
- I removed secrets, tokens, account details, request credentials, and personal data.
- Dominant language
- TypeScript
- Stars
- 16.9k
- Forks
- 1.3k
- Avg merge
- 4h 52m
- Merged PRs (30d)
- 609
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from lidge-jun/opencodex
-
account-pool enhancement proxy
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Maintainers usually reply within 1 day
-
[Bug]: Native Messages lane never logs requestedEffort for Claude Code requests (re-file of #5453)Openbug proxy
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
Maintainers usually reply within 1 day
-
account-pool bug
Difficulty 4/5 3-5 days Newbie friendliness 48/100
lidge-jun/opencodex#6749 · 1 comment ·
Maintainers usually reply within 1 day
-
bug platform
Difficulty 5/5 Over a week Newbie friendliness 30/100
Maintainers usually reply within 1 day
All issues in lidge-jun/opencodex
Similar issues
-
bug priority:low ready-for-dev
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
Automattic/data-liberation-agent#685 ·
Maintainers usually reply within 1 day
-
Business
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
txn2/mcp-data-platform#2063 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 77/100
Crosstalk-Solutions/project-nomad#1427 ·
Maintainers usually reply within 2 days