Custom "anthropic" provider does not enforce provider.max_prompt_tokens — sessions grow past the model context window until a hard 400
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
Direzione di ricerca
Parti dal percorso session-create/resume del provider anthropic personalizzato e traccia come vengono gestiti max_prompt_tokens e le soglie di infinite_sessions prima dell’invio delle richieste. Confronta questo percorso con quello del backend Copilot, dove la compattazione funziona, riproduci la trascrizione crescente e conferma che i prompt rimangano entro il budget configurato e che le sessioni non riuscite possano recuperare.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
When a custom provider of type: "anthropic" is configured with max_prompt_tokens, the SDK does not appear to enforce that budget. A long-running session keeps accumulating transcript until the request exceeds the model's context window and the provider rejects it with HTTP 400 -- after which the session is permanently unusable.
Configuration
The provider is created on every session create/resume with an explicit prompt budget:
{
"type": "anthropic",
"base_url": "...",
"model_id": "claude-sonnet-5",
"max_output_tokens": 32768,
"max_prompt_tokens": 967232
}
max_prompt_tokens is derived as context_window (1,000,000) - max_output_tokens (32,768) = 967,232.
Expected
The SDK compacts (or otherwise bounds the prompt) before crossing max_prompt_tokens = 967232.
Actual
The transcript grew unbounded to 1,001,142 tokens -- 33,910 past the configured budget, and past the model's 1,000,000 hard limit:
400 invalid_request_error
"prompt is too long: 1001142 tokens > 1000000 maximum"
The session had run ~18 successful turns over ~3 hours, with the serialized request growing steadily (~1.98 MB -> ~2.04 MB) before crossing the limit. Tool count was constant throughout, so the growth is accumulated conversation history rather than tool schemas.
Two additional observations
-
infinite_sessionsthresholds also appear inert on this path.background_compaction_threshold/buffer_exhaustion_thresholdare sent on every turn but appear to have no effect for theanthropicprovider (they do take effect on the Copilot backend path). So neither the threshold-based compaction nor themax_prompt_tokensbudget bounded the transcript. -
The session actively degrades after the first failure. Once over the limit, continued turns keep appending to the transcript -- request size grew from ~2.044 MB to ~2.065 MB across ~30 consecutive failed turns. There is no back-off, trim, or compaction triggered by the 400, so the session can never self-recover; every subsequent turn fails immediately (~1.5s vs. the 38s first failure).
Impact
Every turn in an affected session fails permanently. The only recovery is to start a new session, and nothing in the surfaced error indicates that to the user. Because the failure is a deterministic 400, retry suppression correctly kicks in -- but that just means the session is durably wedged.
Environment
- SDK 1.0.7 / Copilot CLI 1.0.71
- Custom
anthropicprovider over an OpenAI-incompatible relay endpoint - Model
claude-sonnet-5(1,000,000-token context window)
Ask
Should provider.max_prompt_tokens be enforced on the anthropic provider path (and/or should infinite_sessions compaction apply there)? If enforcement is intentionally backend-only today, it would help to document that clearly, since the field is accepted without warning and silently has no effect.
- Lingua principale
- TypeScript
- Stelle
- 10.5k
- Fork
- 1.5k
- Merge medio
- 1g 7h
- PR unite (30g)
- 98
Preparare l'ambiente
Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di github/copilot-sdk
-
documentation
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
github/copilot-sdk#2804 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
github/copilot-sdk#2798 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
github/copilot-sdk#2793 ·
I maintainer di solito rispondono entro 1 giorno
-
agentic-workflows
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
github/copilot-sdk#2782 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
github/copilot-sdk#2781 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di github/copilot-sdk
Issue simili
-
Add: PRO TV Chisinau SDApertacheck:failed streams:add
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
iptv-org/iptv#53974 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
interledger/rafiki#3986 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Doist/todoist-cli#576 ·
I maintainer di solito rispondono entro 1 giorno
-
Suggestion: document (or optionally add) a cheaper-model config for find-skills on Claude CodeApertafeature
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
vercel-labs/skills#2370 ·
I maintainer di solito rispondono entro 1 giorno
-
🐛 Bug supabase/cli
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno