Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug]: Non-Codex clients on public endpoints get zero prompt-cache reuse on ChatGPT-backed OpenAI models (no synthesized session_id/originator)

Open
#6,520 3 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
62/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
typescript

Research direction

Start by reviewing the treatment in #6482 and the affinity-key logic in src/oauth/anthropic-routing.ts, then trace the OpenAI/ChatGPT route for missing session_id and originator headers. Done means generic clients receive stable synthesized identity headers and repeated requests show prompt-cache reuse comparable to Codex clients.

Written by the indexing model from the issue text.

Description

bug proxy
Client or integration

Other — a third-party OpenAI/Anthropic-compatible coding agent (ZCode desktop) wired to the public local endpoints, plus plain curl. Not the Grok-managed surface.

Area

Proxy and routing

Version

2.74.0 and 2.76.0 (latest npm at report time; both reproduce)

Operating system

macOS 26 (Darwin 25.6.0, arm64)

Summary

Non-Codex clients that hit the public endpoints (/v1/responses, /v1/messages) get zero prompt-cache reuse on ChatGPT-backed OpenAI models, even with byte-identical prefixes across turns: every request reports input_tokens_details.cached_tokens: 0 and — via usage.attribution — cache_write_tokens: 0, i.e. the backend never even warms the prefix.

Same machine, same single OpenAI (ChatGPT) account, same model (gpt-6.1-sol): real Codex client sessions (which also route through opencodex via openai_base_url = 127.0.0.1:10100/v1) showed up to 226,432 cached_input_tokens in ~/.codex/sessions rollouts the same evening, so neither the account nor the model is the variable.

Cost impact example: one 25-turn agent task through /v1/responses billed 1.92M uncached input tokens in ~13 minutes.

What I ruled out (tested on 2.74.0 and latest npm 2.76.0)
Variation cached_tokens on repeat request
/v1/messages (claude inbound synthesizes a prompt_cache_key via the system+tools cohort fallback) 0
/v1/responses with an explicit body prompt_cache_key 0
Prefix sizes ~1.8K / 3.3K / 7.0K tokens; 5s and 30s gaps 0
stream: true vs non-stream 0
Upgrading 2.74.0 → 2.76.0 0
Same request + originator: Codex Desktop + session_id: <stable id> headers 2176 / 2341 (93%)

Prefix stability was verified by hashing tools/system/first-user across all 25 captured requests (identical), and the account pool has exactly one account.

This matches the root cause confirmed in #6481 ("The Codex Codex backend only reuses a warmed prefix when session_id is present"), but #6482 only forwards a session identity on the managed Grok surface (x-opencodex-grok: 1). Generic clients — third-party agents, Claude-Code-compatible tools, curl — still get no synthesis, so they pay full price on every turn.

Reproduction

Two identical requests, 5 s apart (<PREFIX> ≈ a stable ~3300-token instructions string):

curl -s http://127.0.0.1:10100/v1/responses \
  -H 'content-type: application/json' \
  -H 'authorization: Bearer <loopback-key>' \
  -d '{"model":"gpt-6.1-sol","max_output_tokens":32,
       "instructions":"<PREFIX>",
       "input":[{"role":"user","content":[{"type":"input_text","text":"hi"}]}]}' \
  | jq '.usage.input_tokens_details'   # both times: {"cache_write_tokens":0,"cached_tokens":0}

# now add the two Codex identity headers to the same pair of requests:
#   -H 'originator: Codex Desktop' -H 'session_id: <stable-id>'
# → request 2: cached_tokens ≈ 93–97% of input
Suggested fix

Extend the #6482 treatment beyond the Grok surface: on the OpenAI/ChatGPT route, when the caller sends no session_id header, synthesize a stable one from identity the proxy already computes — e.g. the inbound x-session-id header, or the affinity key from anthropicSessionKeyFromParts in src/oauth/anthropic-routing.ts — and synthesize an originator when absent. That would bring third-party clients to parity with Codex clients.

Workaround in use

A ~40-line local header-injection shim (listens on 127.0.0.1:10101, forwards to :10100, adds originator: Codex Desktop and derives session_id from the client's x-session-id), with the third-party client's baseUrl pointed at the shim. Verified end-to-end: 3968/4101 cached (96.7%) on the second request. Happy to turn the synthesis into a PR if the approach sounds right.

Related: #6481 / #6482 (Grok surface), #4245 (Cursor), #3433 (Hermes).

Dominant language
TypeScript
Stars
16.9k
Forks
1.3k
Avg merge
4h 58m
Merged PRs (30d)
616

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from lidge-jun/opencodex

All issues in lidge-jun/opencodex

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.