[feature] Expose context cache TTL selection (prompt_cache_options.ttl) for the 1h tier
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 55/100
- issue の種類
- 機能追加
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- python, typescript
- 領域
- api, backend-api-design, cli
調査の方向性
prompt_cache_key がすでに生成パラメータに渡されている packages/kosong/src/kosong/chat_provider/kimi.py から始め、provider の設定が Kimi の coding endpoint に届くまでを追跡してください。その endpoint が prompt_cache_options.ttl をサポートしているか、また同梱されている agent-core-v2 bundle が encodeCacheKey をどのように扱うかを確認してください。5m のデフォルトが文書化され、サポートされる場合に 1h を選択でき、それ以外では互換性が保たれ、設定または環境変数による override がカバーされていれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Summary
The Kimi API supports prompt_cache_options.ttl with two tiers — "5m" (default) and "1h" — for context-cache writes. Kimi Code CLI 2.0.0 has no way to select the cache tier: every request rides the default 5m, and there is no config key, env var, or CLI flag to opt into 1h.
Request: expose a provider-level option (e.g. [providers.kimi] cache_ttl = "5m" | "1h", default "5m", optionally overridable per-run via KIMI_CACHE_TTL), and pass it through as prompt_cache_options.ttl on providers/endpoints that support the field.
Motivation and cost math
Cadenced production workflows — scheduled report generation on a fixed 30-minute cadence (48 turns/day), with a long-lived supervisory session holding a stable prefix (system prompt, tool definitions, knowledge content).
Naively, the 1h tier looks pricier: cache writes cost ¥40/M vs ¥20/M for 5m (USD mirrors: $6/M vs $3/M). What makes it win is hit-renewal: after the first write, every hit within the TTL renews it at the cache-read price and no further writes are billed.
Per day, for a ~1M-token stable prefix at a 30-minute cadence:
5mdefault: the gap exceeds the TTL every time → 48 full writes/day → 48 × ¥20 = ¥960/M/day1htier: first turn writes once, the next 47 turns hit at ¥2/M and renew → ¥40 + 47 × ¥2 = ¥134/M/day
≈ 7× cheaper per day for the same workload, plus lower first-token latency on hits. Without this arithmetic, the feature request looks self-defeating ("1h costs 2× more"); the cadence case is where it pays.
Evidence
- The API field exists and is documented (
docs/api/chat, mirrored on platform.moonshot.cn / platform.kimi.ai / platform.moonshot.ai):prompt_cache_options.mode="implicit"(only supported mode; auto-writes the request prefix to cache)prompt_cache_options.ttl="5m" | "1h"— "Lifetime of the written cache. Only 5m and 1h are supported (default 5m); the two tiers are independent."- Behavior notes: cache is isolated at org granularity; a hit within the TTL renews it and is charged only the cache-read price (no re-write fee).
- Pricing pages (CNY and USD mirrors): cache writes billed per tier — 5m tier ¥20/M vs 1h tier ¥40/M (kimi-k3); cache-hit input ¥2/M (CNY) / $0.30/M (USD); regular input ¥20/M / $3/M.
- Current CLI behavior is consistent with the 5m default. Probe method (offered in full below): warm a ≥256-token prompt, then re-probe at intervals (300s / 1800s / 3600s) and compare
usage.prompt_tokens_details.cached_tokens. Intervals past the default tier's lifetime miss; turns under ~5 minutes hit. (Independent corroboration with the same method and numbers: pi-provider-kimi-code caching notes, measured TTL in [300s, 1800s).) - No knob exists in CLI 2.0.0 —
config.tomlhas no cache-related keys; the official env-vars page lists none;tui.tomloffers onlycache_expiry_hint(a reminder toggle, not a TTL control). The CLI emitsprompt_cache_key— but sendingprompt_cache_keydoes not let the caller select a cache tier; there is noprompt_cache_optionsanywhere in the client path (verified in both the open-source package and the shipped 2.0.0 binary, see below).
Proposal
- Add
cache_ttl = "5m" | "1h"under[providers.kimi](default"5m"to preserve current behavior and pricing), passed through asprompt_cache_options.ttlon providers/endpoints that document the field; omit or ignore it elsewhere for compatibility. Kimi Code talks to the coding endpoint whiledocs/api/chatdocuments the chat endpoint — if the coding endpoint does not honor this field, please document that explicitly so client code can gate on it. - Optional env override:
KIMI_CACHE_TTL=5m|1hfor per-run selection (cadenced batch jobs vs interactive sessions). - Surface the pricing implication in the option's docs (1h-tier writes cost 2× the 5m tier; hits renew the TTL and are billed at the cache-read rate), so users choose knowingly.
Code landing spots (for maintainers' convenience)
- Open-source repo:
packages/kosong/src/kosong/chat_provider/kimi.pyalready carries aprompt_cache_keyslot in the generation parameters (currently line ~98) —prompt_cache_optionswould hang naturally at the same site. - Shipped binary (2.0.0, JS bundle, agent-core-v2): cache handling lives in a
kimiOpenAITrait-styleencodeCacheKeythat only ever emits{ prompt_cache_key: key }. - These may or may not be the same codebase — could maintainers clarify which one is production for the 2.0.0 CLI?
Notes
- The
1htier pays off for cadenced workflows with stable prefixes and turn gaps between 5 and 60 minutes; it is a poor fit for continuously-changing contexts (any prefix change — includingsystem/toolschanges — cold-starts a new cache entry). - I can provide a minimal repro script for the current 5m-default behavior (warm/interval-probe with
cached_tokensevidence) on request.
- 主要言語
- TypeScript
- スター
- 7.7k
- フォーク
- 1.3k
- 平均マージ
- 12時間 35分
- マージ済み PR(30日)
- 311
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
MoonshotAI/kimi-code のほかの issue
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
MoonshotAI/kimi-code#4040 ·
メンテナーはふだん 1 日以内に返信
-
顶栏字体不随字体大小缩放bugオープンbug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
MoonshotAI/kimi-code#4039 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
MoonshotAI/kimi-code#4010 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
MoonshotAI/kimi-code#4008 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
MoonshotAI/kimi-code#3947 ·
メンテナーはふだん 1 日以内に返信
MoonshotAI/kimi-code の issue をすべて見る
似ている issue
-
needs:triage
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信
-
ai-discovered
難易度 2/5 1〜3時間 初心者へのやさしさ 83/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
jessepollak/home#1627 ·
メンテナーはふだん 1 日以内に返信
-
agent-canvas bug llm priority:low ready-for-dev
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
OpenHands/OpenHands#17806 · コメント 3 件 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
radius-project/ai-extensions#923 ·
メンテナーはふだん 1 日以内に返信