applyCaching gate keyed on model-name substring, not capability — non-Anthropic cacheable models (openrouter/openai-compatible/copilot) never get cache_control
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 45/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- typescript
Hướng nghiên cứu
Bắt đầu trong packages/opencode/src/provider/transform.ts, đặc biệt là ProviderTransform.message() ở các dòng 285-297 và applyCaching() ở các dòng 196-210, sau đó kiểm tra schema capabilities trong provider.ts:787. So sánh gate với các tùy chọn provider được hỗ trợ và phạm vi coverage hiện có; hoàn tất khi capability hoặc quy tắc provider được chọn được định nghĩa rõ ràng và các model có thể cache không phải Anthropic nhận được các directive dự kiến mà không làm suy giảm hành vi hiện có.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Description
Found via static analysis of prompt-cache anti-patterns (CacheLint), then confirmed by hand on main @ f0fb1e1.
ProviderTransform.message() decides whether to call applyCaching() using a hard-coded model-name gate (packages/opencode/src/provider/transform.ts:285-297):
if (
(model.providerID === "anthropic" ||
model.providerID === "google-vertex-anthropic" ||
model.providerID === "altimate-backend" ||
model.api.id.includes("anthropic") ||
model.api.id.includes("claude") ||
model.id.includes("anthropic") ||
model.id.includes("claude") ||
model.api.npm === "@ai-sdk/anthropic") &&
model.api.npm !== "@ai-sdk/gateway"
) {
msgs = applyCaching(msgs, model)
}
But applyCaching() itself already defines cache directives for five providers, not just Anthropic (transform.ts:196-210):
const providerOptions = {
anthropic: { cacheControl: { type: "ephemeral" } },
openrouter: { cacheControl: { type: "ephemeral" } },
bedrock: { cachePoint: { type: "default" } },
openaiCompatible: { cache_control: { type: "ephemeral" } },
copilot: { copilot_cache_control:{ type: "ephemeral" } },
}
So a cacheable model served through openrouter / openai-compatible / copilot whose id contains neither claude nor anthropic (e.g. a GPT / Gemini / Qwen / Kimi routed through those providers) never enters applyCaching() and never gets a cache breakpoint — even though the function clearly intends to cache it.
This is an under-claim: caching silently fails to engage. It is not cache-busting, and Anthropic-named models are unaffected.
Impact
For an affected model, the system prefix (system prompt + tool schema + earlier turns) is re-sent at full input price on every turn instead of being read from cache. On long agentic loops that is roughly the usual cached-prefix discount forgone each turn, plus higher TTFB — for exactly the self-hosted / BYO-LLM users the project targets. The blast radius is bounded to non-Anthropic-named models that genuinely support explicit cache_control via openrouter/openai-compatible/copilot.
Steps to reproduce
- Configure a cacheable model through
openrouter(oropenai-compatible/copilot) whose id does not containclaude/anthropic(e.g. an OpenRouter-served model that honorscache_control). - Run a multi-turn session.
- Observe that no
cacheControl/cache_controlprovider option is stamped on the system/last-user blocks (theapplyCachingbranch is skipped), so the prefix is billed as fresh input every turn.
Suggested fix (for discussion)
Decouple the applyCaching gate from the model-name list and drive it off an explicit capability/provider-support signal, so the gate matches the set of providers applyCaching already knows how to cache:
- introduce a
capabilities.cachingflag on the model (thecapabilitiesschema inprovider.ts:787currently has no caching field), populated from the model registry; or - gate on a per-provider "supports prompt cache" set covering
anthropic / openrouter / bedrock / openaiCompatible / copilot(the same five keys already inproviderOptions).
I want to flag the design angle rather than send a drive-by PR: this is a hot path, and #891 was deliberately deferred for the same "needs design + careful testing, don't regress gateway cache-hit rates" reason. It also overlaps with #891's goal of having a single source of truth for the cache-control gate. I'm happy to open a PR if a maintainer confirms the preferred shape (capability flag vs provider set) and that emitting these directives for non-Anthropic providers is intended.
Caveat: confirmed present and unguarded on main @ f0fb1e1; line numbers may drift.
- Ngôn ngữ chính
- TypeScript
- Star
- 813
- Fork
- 134
- Merge trung bình
- 2 ngày 3 giờ
- Pull request đã merge (30 ngày)
- 65
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của AltimateAI/altimate-code
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
AltimateAI/altimate-code#1359 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
AltimateAI/altimate-code#1323 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
AltimateAI/altimate-code#1288 ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
AltimateAI/altimate-code#1285 ·
-
privacy: Altimate Base consent dialog no longer discloses persistent per-installation identifier Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
AltimateAI/altimate-code#1284 ·
Tất cả issue của AltimateAI/altimate-code
Issue tương tự
-
bug(cli): hapi doctor inline-media prints a fabricated B:\ helper-script path in packaged installs Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
-
Crush Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 85/100
catppuccin/catppuccin#3125 ·
-
Add a SECURITY.md Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
ElementsProject/cln-application#167 · 1 bình luận · 1 reaction ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
Quantco/pnpm-licenses#17 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100