Context caching (ContextCacheConfig) is a silent no-op: config is plumbed to InvocationContext but never read
I maintainer di solito rispondono entro 1 giorno
@hemasekhar-p ci sta già lavorando.
Dal 23/9/2026.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
Direzione di ricerca
Start by tracing ContextCacheConfig.java through App.java, Runner.java, and InvocationContext.java, then run the supplied recording-BaseLlm reproduction to confirm the request is unchanged. Compare the relevant adk-python cache manager and flow files before deciding the Java package structure. Done requires either working cache creation and request metadata with tests, or an explicit loud failure and corrected support documentation, depending on maintainer direction.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
ContextCacheConfig is accepted by App.Builder and plumbed all the way to InvocationContext, but nothing in the library ever reads it. Setting it is a silent no-op: no cached content is created, no cachedContent is set on the request, and the full system instruction is sent as tokens on every call.
The context caching docs list Java as supported ("Python v1.15.0+, Java v0.1.0+, Kotlin v0.7.0+"), and RFC #543 was closed with "this feature was successfully implemented in version 0.6.0". I believe that closure was made in error: the commit it points to (12defeed, PR #822) added only the configuration class, and its own message says so — "introduces context caching configuration for apps" (+59 lines in ContextCacheConfig.java, +19/-2 in App.java). PR #823 then added the InvocationContext plumbing. No implementation PR followed.
Because the config is accepted, documented as supported, and plumbed to the point of use, there is no signal that it does nothing — you only find out from a cache-hit rate that never moves.
The chain, traced on main
App.Builder.contextCacheConfig(cfg)stores it onApp.Runner.Builder.build()readsapp.contextCacheConfig()(Runner.java:176) and assigns theRunnerfield.Runner.newInvocationContextBuilder()passes it toInvocationContext.Builder.contextCacheConfig(...).InvocationContext.Builder(InvocationContext)copies it, so child contexts carry it.- Reads of that field:
equals/hashCode, and the public getterInvocationContext.contextCacheConfig()(InvocationContext.java:303). - That getter has no callers. The trail ends there.
Supporting checks on main:
- Zero occurrences of
CreateCachedContentConfig,CachedContent,caches.createorCacheMetadataanywhere in the repo. ContextCacheConfig.getTtlString()— javadoc'd "Returns TTL as string format for cache creation", i.e. written specifically for this feature — has no callers either.- The only cache-adjacent code reads
usageMetadata.cachedContentTokenCountback off responses for reporting (GeminiUtil,BigQueryAgentAnalyticsPlugin), which is a different thing. - Verified identically against the published 1.9.0 and 1.10.1 sources jars, so this is not new.
Reproduction
Configure caching on an App, run one turn against a recording BaseLlm, and inspect the request:
ContextCacheConfig caching =
new ContextCacheConfig(/* maxInvocations= */ 10, Duration.ofHours(1), /* minTokens= */ 0);
App app = App.builder()
.name("repro")
.rootAgent(LlmAgent.builder().name("repro").model(recordingLlm)
.instruction("...32KB of static instruction...").build())
.contextCacheConfig(caching)
.build();
Runner runner = Runner.builder()
.app(app)
.sessionService(new InMemorySessionService())
.artifactService(new InMemoryArtifactService())
.build();
// run one turn, then on the LlmRequest the model received:
assertThat(request.config().get().cachedContent()).isEmpty(); // passes
assertThat(request.config().get().systemInstruction()).isPresent(); // passes — full text, inline
assertThat(app.contextCacheConfig()).isEqualTo(caching); // passes — config was kept
Observed, with minTokens = 0 and a ~32KB instruction so no size floor applies:
| Prompt supplied as | cachedContent on the request |
System instruction |
|---|---|---|
Instruction.Provider |
empty | inline, full text |
static .instruction(String) |
empty | inline, full text |
systemInstruction on generateContentConfig |
empty | inline, full text |
The request is byte-identical with the config present and absent. Response-side too: a response reporting no cachedContentTokenCount reaches the event with none invented for it, and one reporting cachedContentTokenCount passes through untouched (BaseLlmFlow copies usageMetadata verbatim) — so the reporting channel works and the absence above is genuine.
Impact
For an agent with a large static instruction this is the whole cost saving silently not happening. In our case the evaluator prompt is ~7k tokens, about two thirds of every request's input tokens. Implicit caching does not substitute for it at low request rates — entries live a few minutes, so at single-digit requests per minute spread over several prompt variants we measured roughly a 1% hit rate.
Suggested resolution
Either would resolve the confusion; the first is obviously preferable:
- Implement it, porting
gemini_context_cache_manager.py,cache_metadata.pyandflows/llm_flows/context/_cache.pyfrom adk-python. I'm happy to attempt this as a PR if that would be welcome, but given the size (~1k lines of Python, with non-trivial fingerprint/invalidation state) and that this is a design-sensitive area, I'd rather check first whether it's already planned internally or whether there's a preferred shape — the package-structure question from #543 was never settled publicly. - Failing that, make the no-op loud — reopen #543, correct the docs' Java support claim, and either javadoc
ContextCacheConfigas not-yet-implemented on Java or log a warning when anAppis built with one. That costs little and stops the next person spending a day proving the negative.
Happy to supply the full repro as a test case if useful.
Environment
- adk-java
main, plus published 1.9.0 and 1.10.1 (behaviour identical) - Java 25,
google-genai1.58.0
- Lingua principale
- Java
- Stelle
- 1.7k
- Fork
- 421
- Merge medio
- 3g 9h
- PR unite (30g)
- 29
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di google/adk-java
-
BaseLlmFlow nests each step inside the previous one and overflows the stack after a few hundred LLM callsForse già presa @hemasekhar-p l’ha presa 2 giorni fa. Apertaneeds review
Difficoltà 4/5 3-5 giorni Idoneità per principianti 68/100
google/adk-java#1564 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
AgentTool runs the wrapped agent with the default RunConfig instead of the caller'sForse già presa @hemasekhar-p l’ha presa 2 giorni fa. Apertaneeds review
Difficoltà 3/5 1-2 giorni Idoneità per principianti 72/100
google/adk-java#1562 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Approved tool call re-runs on every later user turn if it never got a function responseForse già presa @hemasekhar-p l’ha presa 5 giorni fa. Apertaneeds review
google/adk-java#1556 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
LocalSkillSource.listResources returns backslash-separated paths on WindowsForse già presa @hemasekhar-p l’ha presa 6 giorni fa. Apertaneeds review
google/adk-java#1541 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
MCP toolset: in-model built-ins (e.g. google_search) can be shadowed by a server tool; a server tool named set_model_response aborts the runForse già presa @hemasekhar-p l’ha presa 14 giorni fa. Apertaneeds review
google/adk-java#1513 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di google/adk-java
Issue simili
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
apache/arrow-java#1311 ·
I maintainer di solito rispondono entro 2 giorni
-
bug triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno
-
security
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
IBM/networking-java-sdk#204 ·
-
bug Technical Debt
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
avniproject/avni-server#1080 ·