Custom "anthropic" provider does not enforce provider.max_prompt_tokens — sessions grow past the model context window until a hard 400
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 48/100
Hướng nghiên cứu
Bắt đầu từ đường dẫn session-create/resume của provider anthropic tùy chỉnh và theo dõi cách max_prompt_tokens cùng các ngưỡng infinite_sessions được xử lý trước khi các request được gửi đi. So sánh đường dẫn đó với đường dẫn backend Copilot, nơi tính năng compaction hoạt động, tái hiện transcript ngày càng dài và xác nhận rằng các prompt vẫn nằm trong ngân sách đã cấu hình तथा các session bị lỗi có thể khôi phục.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
When a custom provider of type: "anthropic" is configured with max_prompt_tokens, the SDK does not appear to enforce that budget. A long-running session keeps accumulating transcript until the request exceeds the model's context window and the provider rejects it with HTTP 400 -- after which the session is permanently unusable.
Configuration
The provider is created on every session create/resume with an explicit prompt budget:
{
"type": "anthropic",
"base_url": "...",
"model_id": "claude-sonnet-5",
"max_output_tokens": 32768,
"max_prompt_tokens": 967232
}
max_prompt_tokens is derived as context_window (1,000,000) - max_output_tokens (32,768) = 967,232.
Expected
The SDK compacts (or otherwise bounds the prompt) before crossing max_prompt_tokens = 967232.
Actual
The transcript grew unbounded to 1,001,142 tokens -- 33,910 past the configured budget, and past the model's 1,000,000 hard limit:
400 invalid_request_error
"prompt is too long: 1001142 tokens > 1000000 maximum"
The session had run ~18 successful turns over ~3 hours, with the serialized request growing steadily (~1.98 MB -> ~2.04 MB) before crossing the limit. Tool count was constant throughout, so the growth is accumulated conversation history rather than tool schemas.
Two additional observations
-
infinite_sessionsthresholds also appear inert on this path.background_compaction_threshold/buffer_exhaustion_thresholdare sent on every turn but appear to have no effect for theanthropicprovider (they do take effect on the Copilot backend path). So neither the threshold-based compaction nor themax_prompt_tokensbudget bounded the transcript. -
The session actively degrades after the first failure. Once over the limit, continued turns keep appending to the transcript -- request size grew from ~2.044 MB to ~2.065 MB across ~30 consecutive failed turns. There is no back-off, trim, or compaction triggered by the 400, so the session can never self-recover; every subsequent turn fails immediately (~1.5s vs. the 38s first failure).
Impact
Every turn in an affected session fails permanently. The only recovery is to start a new session, and nothing in the surfaced error indicates that to the user. Because the failure is a deterministic 400, retry suppression correctly kicks in -- but that just means the session is durably wedged.
Environment
- SDK 1.0.7 / Copilot CLI 1.0.71
- Custom
anthropicprovider over an OpenAI-incompatible relay endpoint - Model
claude-sonnet-5(1,000,000-token context window)
Ask
Should provider.max_prompt_tokens be enforced on the anthropic provider path (and/or should infinite_sessions compaction apply there)? If enforcement is intentionally backend-only today, it would help to document that clearly, since the field is accepted without warning and silently has no effect.
- Ngôn ngữ chính
- TypeScript
- Star
- 10.5k
- Fork
- 1.5k
- Merge trung bình
- 1 ngày 7 giờ
- Pull request đã merge (30 ngày)
- 98
Chuẩn bị môi trường
Khởi chạy dev container của dự án ngay trên trình duyệt, bằng tài khoản GitHub của bạn.
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của github/copilot-sdk
-
documentation
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
github/copilot-sdk#2804 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
github/copilot-sdk#2798 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
github/copilot-sdk#2793 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
agentic-workflows
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
github/copilot-sdk#2782 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
github/copilot-sdk#2781 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của github/copilot-sdk
Issue tương tự
-
enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
tomjn/coilbox-hub#454 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày
-
api: spanner
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
googleapis/google-cloud-node#9513 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 92/100
Maintainer thường phản hồi trong vòng 1 ngày
-
SegmentedControl calls Math.random() during render, breaking Next.js cacheComponents prerenderingĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
mantinedev/mantine#9244 ·
Maintainer thường phản hồi trong vòng 8 ngày