[Bug]: Non-Codex clients on public endpoints get zero prompt-cache reuse on ChatGPT-backed OpenAI models (no synthesized session_id/originator)
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 62/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- typescript
- Domain
- api, backend-api-design
Research direction
Start by reviewing the treatment in #6482 and the affinity-key logic in src/oauth/anthropic-routing.ts, then trace the OpenAI/ChatGPT route for missing session_id and originator headers. Done means generic clients receive stable synthesized identity headers and repeated requests show prompt-cache reuse comparable to Codex clients.
Written by the indexing model from the issue text.
Description
Client or integration
Other — a third-party OpenAI/Anthropic-compatible coding agent (ZCode desktop) wired to the public local endpoints, plus plain curl. Not the Grok-managed surface.
Area
Proxy and routing
Version
2.74.0 and 2.76.0 (latest npm at report time; both reproduce)
Operating system
macOS 26 (Darwin 25.6.0, arm64)
Summary
Non-Codex clients that hit the public endpoints (/v1/responses, /v1/messages) get zero prompt-cache reuse on ChatGPT-backed OpenAI models, even with byte-identical prefixes across turns: every request reports input_tokens_details.cached_tokens: 0 and — via usage.attribution — cache_write_tokens: 0, i.e. the backend never even warms the prefix.
Same machine, same single OpenAI (ChatGPT) account, same model (gpt-6.1-sol): real Codex client sessions (which also route through opencodex via openai_base_url = 127.0.0.1:10100/v1) showed up to 226,432 cached_input_tokens in ~/.codex/sessions rollouts the same evening, so neither the account nor the model is the variable.
Cost impact example: one 25-turn agent task through /v1/responses billed 1.92M uncached input tokens in ~13 minutes.
What I ruled out (tested on 2.74.0 and latest npm 2.76.0)
| Variation | cached_tokens on repeat request |
|---|---|
/v1/messages (claude inbound synthesizes a prompt_cache_key via the system+tools cohort fallback) |
0 |
/v1/responses with an explicit body prompt_cache_key |
0 |
| Prefix sizes ~1.8K / 3.3K / 7.0K tokens; 5s and 30s gaps | 0 |
stream: true vs non-stream |
0 |
| Upgrading 2.74.0 → 2.76.0 | 0 |
Same request + originator: Codex Desktop + session_id: <stable id> headers |
2176 / 2341 (93%) |
Prefix stability was verified by hashing tools/system/first-user across all 25 captured requests (identical), and the account pool has exactly one account.
This matches the root cause confirmed in #6481 ("The Codex Codex backend only reuses a warmed prefix when session_id is present"), but #6482 only forwards a session identity on the managed Grok surface (x-opencodex-grok: 1). Generic clients — third-party agents, Claude-Code-compatible tools, curl — still get no synthesis, so they pay full price on every turn.
Reproduction
Two identical requests, 5 s apart (<PREFIX> ≈ a stable ~3300-token instructions string):
curl -s http://127.0.0.1:10100/v1/responses \
-H 'content-type: application/json' \
-H 'authorization: Bearer <loopback-key>' \
-d '{"model":"gpt-6.1-sol","max_output_tokens":32,
"instructions":"<PREFIX>",
"input":[{"role":"user","content":[{"type":"input_text","text":"hi"}]}]}' \
| jq '.usage.input_tokens_details' # both times: {"cache_write_tokens":0,"cached_tokens":0}
# now add the two Codex identity headers to the same pair of requests:
# -H 'originator: Codex Desktop' -H 'session_id: <stable-id>'
# → request 2: cached_tokens ≈ 93–97% of input
Suggested fix
Extend the #6482 treatment beyond the Grok surface: on the OpenAI/ChatGPT route, when the caller sends no session_id header, synthesize a stable one from identity the proxy already computes — e.g. the inbound x-session-id header, or the affinity key from anthropicSessionKeyFromParts in src/oauth/anthropic-routing.ts — and synthesize an originator when absent. That would bring third-party clients to parity with Codex clients.
Workaround in use
A ~40-line local header-injection shim (listens on 127.0.0.1:10101, forwards to :10100, adds originator: Codex Desktop and derives session_id from the client's x-session-id), with the third-party client's baseUrl pointed at the shim. Verified end-to-end: 3968/4101 cached (96.7%) on the second request. Happy to turn the synthesis into a PR if the approach sounds right.
Related: #6481 / #6482 (Grok surface), #4245 (Cursor), #3433 (Hermes).
- Dominant language
- TypeScript
- Stars
- 16.9k
- Forks
- 1.3k
- Avg merge
- 4h 58m
- Merged PRs (30d)
- 616
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from lidge-jun/opencodex
-
account-pool enhancement proxy
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Maintainers usually reply within 1 day
-
Auth: paste Service Account JSON(s) as the API key — keychain-backed markers + key pool rotation (re: #6821)Possibly taken @GSL-R claimed this today. Openaccount-pool catalog enhancement
Difficulty 5/5 Over a week Newbie friendliness 22/100
lidge-jun/opencodex#6842 · 1 comment ·
Maintainers usually reply within 1 day
-
account-pool enhancement gui proxy
Difficulty 5/5 Over a week Newbie friendliness 8/100
Maintainers usually reply within 1 day
-
bug platform service
Difficulty 4/5 3-5 days Newbie friendliness 22/100
Maintainers usually reply within 1 day
-
catalog enhancement
Difficulty 4/5 3-5 days Newbie friendliness 35/100
lidge-jun/opencodex#6784 · 2 comments ·
Maintainers usually reply within 1 day
All issues in lidge-jun/opencodex
Similar issues
-
clawsweeper:fix-shape-clear clawsweeper:queueable-fix clawsweeper:source-repro impact:other issue-rating: 🦞 diamond lobster no-stale P2
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
openclaw/openclaw#168089 · 2 comments · 1 reaction ·
Maintainers usually reply within 1 day
-
✨ enhancement needs-discussion
Difficulty 1/5 Under an hour Newbie friendliness 85/100
-
[Bug]: [MCP/CLI] Bare loopback IP addresses (127.0.0.1:port) and hosts with ports fail to navigate due to erroneous scheme inferencePossibly taken @alok-108 claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
microsoft/playwright#43263 ·
Maintainers usually reply within 1 day
-
area:studio type:security
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day
-
enhancement good first issue Stellar Wave trivial
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
StellarCanary/ProtocolCanary-Action#331 ·
Maintainers usually reply within 1 day