Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug] The restart drain fence refuses new requests for the full 60s while Codex exhausts its retries in about six seconds

Closed Beginner friendly
#6,643 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
70/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
typescript
Domain
backend

Research direction

Start in src/server/management/system-restart.ts at the deadline calculation around line 366 and the drain call around line 393. Check how the restart handoff uses the remaining deadline, then verify that bounding the refusal window still lets the drain and replacement startup proceed. Done means new requests are refused for no longer than the handoff cap and a client can retry against the replacement.

Written by the indexing model from the issue text.

Description

bug service
Client or integration

Codex CLI

Area

Service lifecycle

Summary

The restart handoff drains for up to MEMORY_DRAIN_RESTART_MS (60s). The intent is to let an in-flight turn finish, but during that whole window every NEW data-plane request is refused, and Codex spends its five sampling retries in roughly six seconds. So the client gives up well before the replacement is listening, and the turn fails even though the restart itself is quick.

Measured on this machine (2026-09-26): with the 60s default, a 42-second fence failed the turn with unexpected status 503 - and that was already after the refusal stopped being a fatal capacity error (see the related issue below). A two-second cap, plus the replacement's roughly one-second bind gap, passed. In other words the fence length, not the refusal itself, is what the client runs out of.

Reproduction
  1. ocx start --port 10100.
  2. Start a Codex turn.
  3. Trigger the drain-and-restart path (dashboard POST /api/system/restart, or ocx service restart) while the turn is in flight.
  4. Watch the turn: the client retries against a listener that is refusing for the full fence, exhausts its retries, and ends with unexpected status 503 instead of continuing against the replacement.
Version

2.78.0

Operating system

macOS 27.0.1 (arm64). The fence is shared code, so this should apply on every platform.

Provider and model

Not provider-specific.

Logs or error output

Source location: src/server/management/system-restart.ts - the deadline is set at line 366 (const restartDeadlineMs = now() + MEMORY_DRAIN_RESTART_MS;) and consumed at line 393 (const remainingMs = Math.max(0, restartDeadlineMs - now());), which is then handed to drain(undefined, remainingMs).

Suggested change

Bound the restart fence rather than the drain - keep the drain free to finish an in-flight turn, but stop refusing new requests for longer than the client can retry:

const RESTART_HANDOFF_INFLIGHT_WAIT_MS = 2_000;
const remainingMs = Math.min(
  Math.max(0, restartDeadlineMs - now()),
  RESTART_HANDOFF_INFLIGHT_WAIT_MS,
);

A turn cut by this cap is retried by the client against the replacement that the same handoff starts, which is why two seconds is enough: the gap it has to cover is the replacement's bind, not the full drain.

This is what we run as a local patch. Happy to open a PR if you would rather review it that way.

Related: the code carried by the refusal is the other half of this handoff problem, filed separately as #6642.

Screenshots and supporting files

None.

Redacted configuration
{ "hostname": "127.0.0.1", "port": 10100 }
Checks
  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.
Dominant language
TypeScript
Stars
16.9k
Forks
1.3k
Avg merge
4h 52m
Merged PRs (30d)
609

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from lidge-jun/opencodex

All issues in lidge-jun/opencodex

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.