Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

[feature] Expose context cache TTL selection (prompt_cache_options.ttl) for the 1h tier

Aperta
#3,942 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
55/100
Tipo di issue
Funzionalità
Chiarezza
Abbastanza chiara
Stato di attività
Attiva
Stack tecnologico
python, typescript

Direzione di ricerca

Inizia da packages/kosong/src/kosong/chat_provider/kimi.py, dove prompt_cache_key è già incluso nei parametri di generazione, e segui il percorso con cui la configurazione del provider raggiunge l’endpoint di coding di Kimi. Verifica se l’endpoint supporta prompt_cache_options.ttl e come il bundle agent-core-v2 incluso gestisce encodeCacheKey. Il lavoro è completo quando sono presenti un valore predefinito documentato di 5m, un comportamento selezionabile di 1h dove supportato, la compatibilità altrove e la copertura per la configurazione o l’override tramite ambiente.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Summary

The Kimi API supports prompt_cache_options.ttl with two tiers — "5m" (default) and "1h" — for context-cache writes. Kimi Code CLI 2.0.0 has no way to select the cache tier: every request rides the default 5m, and there is no config key, env var, or CLI flag to opt into 1h.

Request: expose a provider-level option (e.g. [providers.kimi] cache_ttl = "5m" | "1h", default "5m", optionally overridable per-run via KIMI_CACHE_TTL), and pass it through as prompt_cache_options.ttl on providers/endpoints that support the field.

Motivation and cost math

Cadenced production workflows — scheduled report generation on a fixed 30-minute cadence (48 turns/day), with a long-lived supervisory session holding a stable prefix (system prompt, tool definitions, knowledge content).

Naively, the 1h tier looks pricier: cache writes cost ¥40/M vs ¥20/M for 5m (USD mirrors: $6/M vs $3/M). What makes it win is hit-renewal: after the first write, every hit within the TTL renews it at the cache-read price and no further writes are billed.

Per day, for a ~1M-token stable prefix at a 30-minute cadence:

  • 5m default: the gap exceeds the TTL every time → 48 full writes/day → 48 × ¥20 = ¥960/M/day
  • 1h tier: first turn writes once, the next 47 turns hit at ¥2/M and renew → ¥40 + 47 × ¥2 = ¥134/M/day

≈ 7× cheaper per day for the same workload, plus lower first-token latency on hits. Without this arithmetic, the feature request looks self-defeating ("1h costs 2× more"); the cadence case is where it pays.

Evidence

  1. The API field exists and is documented (docs/api/chat, mirrored on platform.moonshot.cn / platform.kimi.ai / platform.moonshot.ai):
    • prompt_cache_options.mode = "implicit" (only supported mode; auto-writes the request prefix to cache)
    • prompt_cache_options.ttl = "5m" | "1h" — "Lifetime of the written cache. Only 5m and 1h are supported (default 5m); the two tiers are independent."
    • Behavior notes: cache is isolated at org granularity; a hit within the TTL renews it and is charged only the cache-read price (no re-write fee).
  2. Pricing pages (CNY and USD mirrors): cache writes billed per tier — 5m tier ¥20/M vs 1h tier ¥40/M (kimi-k3); cache-hit input ¥2/M (CNY) / $0.30/M (USD); regular input ¥20/M / $3/M.
  3. Current CLI behavior is consistent with the 5m default. Probe method (offered in full below): warm a ≥256-token prompt, then re-probe at intervals (300s / 1800s / 3600s) and compare usage.prompt_tokens_details.cached_tokens. Intervals past the default tier's lifetime miss; turns under ~5 minutes hit. (Independent corroboration with the same method and numbers: pi-provider-kimi-code caching notes, measured TTL in [300s, 1800s).)
  4. No knob exists in CLI 2.0.0 — config.toml has no cache-related keys; the official env-vars page lists none; tui.toml offers only cache_expiry_hint (a reminder toggle, not a TTL control). The CLI emits prompt_cache_key — but sending prompt_cache_key does not let the caller select a cache tier; there is no prompt_cache_options anywhere in the client path (verified in both the open-source package and the shipped 2.0.0 binary, see below).

Proposal

  • Add cache_ttl = "5m" | "1h" under [providers.kimi] (default "5m" to preserve current behavior and pricing), passed through as prompt_cache_options.ttl on providers/endpoints that document the field; omit or ignore it elsewhere for compatibility. Kimi Code talks to the coding endpoint while docs/api/chat documents the chat endpoint — if the coding endpoint does not honor this field, please document that explicitly so client code can gate on it.
  • Optional env override: KIMI_CACHE_TTL=5m|1h for per-run selection (cadenced batch jobs vs interactive sessions).
  • Surface the pricing implication in the option's docs (1h-tier writes cost 2× the 5m tier; hits renew the TTL and are billed at the cache-read rate), so users choose knowingly.

Code landing spots (for maintainers' convenience)

  • Open-source repo: packages/kosong/src/kosong/chat_provider/kimi.py already carries a prompt_cache_key slot in the generation parameters (currently line ~98) — prompt_cache_options would hang naturally at the same site.
  • Shipped binary (2.0.0, JS bundle, agent-core-v2): cache handling lives in a kimiOpenAITrait-style encodeCacheKey that only ever emits { prompt_cache_key: key }.
  • These may or may not be the same codebase — could maintainers clarify which one is production for the 2.0.0 CLI?

Notes

  • The 1h tier pays off for cadenced workflows with stable prefixes and turn gaps between 5 and 60 minutes; it is a poor fit for continuously-changing contexts (any prefix change — including system/tools changes — cold-starts a new cache entry).
  • I can provide a minimal repro script for the current 5m-default behavior (warm/interval-probe with cached_tokens evidence) on request.
Lingua principale
TypeScript
Stelle
7.7k
Fork
1.3k
Merge medio
12h 57m
PR unite (30g)
308

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di MoonshotAI/kimi-code

Tutte le issue di MoonshotAI/kimi-code

Issue simili

Altre issue su TypeScript

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.