Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug] Windows response spill ECAPACITY at 1 GiB yields continuation tombstones while memory remains healthy (2.80.0)

Closed
#6,747 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
30/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
typescript
Domain
backend

Research direction

Start by reading src/responses/state.ts and the ECAPACITY fail-closed paths in src/responses/state/spill-queue.ts; then trace how the memory and spill counters are exposed through /api/system/memory. The issue calls for design work: keep the spill storage bounded while handling expired or superseded entries safely, and make failures for live continuations clearly diagnosable with an actionable recovery path.

Written by the indexing model from the issue text.

Description

bug platform
Client or integration

Codex App

Area

Platform (Windows / macOS / Linux)

Summary

On a busy Windows 11 OpenCodex 2.80.0 service, the durable Responses continuation spill storage repeatedly reports ECAPACITY as its ~1 GiB aggregate cap fills. The proxy /healthz and memory status remain healthy; cumulative spillWriteFailures and tombstoneCount grow. This threatens the reliable replay/resume of active threads even though the system still has tens of GiB of free disk space. This is not a request to uncap the spill directory: the hard cap is a valuable invariant. We need a bounded, safe eviction/admission/recovery strategy or a clear actionable "continuation not durable" health/diagnostic signal.

Reproduction
  1. Run official OpenCodex 2.80.0 on Windows 11 with approximately 7–10 concurrent legitimate Codex requests/threads for several hours. Allow the ordinary production Responses state store and spill queue to manage continuation payloads; do not override internal test constants.
  2. Periodically read supported /api/system/memory via local authenticated management. Observe resident/stub/tombstone counts, spill bytes and write failures.
  3. After spill payloads approach 1 GiB, spillWriteFailures increases while spillLastWriteFailureCode=ECAPACITY. Health and memory status remain "healthy" after a later successful write (spillWriteConsecutiveFailures=0).
  4. The in-memory response state cap remains 1000 entries and the durable aggregate is intentionally bounded by MAX_SPILLED_RESPONSE_BYTES = 1024 * 1024 * 1024. See src/responses/state.ts and the ECAPACITY fail-closed paths in src/responses/state/spill-queue.ts.

This case differs from #5369's unbounded disk growth and #3522's Windows ACL-related stall. Here the bounded cap is enforced, but normal continued throughput still creates many refused continuation writes/tombstones.

Version

2.80.0 (official npm published package, Bun 1.4.0).

Operating system

Windows 11 desktop.

Provider and model

Canonical openai ChatGPT-login pooled route, gpt-6.1-sol across multiple reasoning strengths; exact model selection is not needed to reproduce the spill capacity problem.

Logs or error output
2026-10-07 12:17 (local UTC+08:00): spillWriteFailures=0; tombstoneCount=0
2026-10-08 12:17: spillWriteFailures=84; tombstoneCount=64
2026-10-08 13:17: spillWriteFailures=220; tombstoneCount=5; spillPayloadBytes≈913MiB
2026-10-08 13:32: spillWriteFailures=258; tombstoneCount=43; spillPayloadBytes≈994MiB
2026-10-08 13:47: spillWriteFailures=317; tombstoneCount=97; spillPayloadBytes≈1008MiB

Additional isolated snapshot during active service:
spillWriteStatus=healthy
spillWriteConsecutiveFailures=0
spillLastWriteFailureCode=ECAPACITY
spillWriteFailures=197 at sampling time
spillReadFailures=0

Memory monitor: app-owned budget overrun=0 (healthy)
Disk: >50GiB free on system drive
Production service: /healthz=200 and steady PID

These values are cumulative counters from one process, not absolute counts of unique failed user conversations. No claim is made that ECAPACITY caused the concurrent upstream response.failed 502s (which have separate telemetry).

Screenshots and supporting files

Only aggregate, redacted numeric telemetry is included. Detailed non-sensitive memory snapshots can be shared on request; no conversation payloads or spill files.

Redacted configuration
{
  "providers": {"openai": {"adapter": "openai-responses", "authMode": "forward", "codexAccountMode": "pool"}},
  "websockets": false,
  "appOwnedMemoryBudgetMb": 256
}

The storage cap is the compiled default, not a user override. We have not attempted to increase it or delete active spill files.

Desired outcome: Bound disk usage and preserve durable continuation replay for still-live threads where possible; evict truly expired/superseded references first, report failure status clearly, and expose a supported recovery path when the live set can no longer fit.

Checks
  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.
Dominant language
TypeScript
Stars
16.9k
Forks
1.3k
Avg merge
4h 58m
Merged PRs (30d)
616

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from lidge-jun/opencodex

All issues in lidge-jun/opencodex

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.