Context caching (ContextCacheConfig) is a silent no-op: config is plumbed to InvocationContext but never read
Maintainers usually reply within 1 day
@hemasekhar-p is already working on this.
Since Sep 23, 2026.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
Research direction
Start by tracing ContextCacheConfig.java through App.java, Runner.java, and InvocationContext.java, then run the supplied recording-BaseLlm reproduction to confirm the request is unchanged. Compare the relevant adk-python cache manager and flow files before deciding the Java package structure. Done requires either working cache creation and request metadata with tests, or an explicit loud failure and corrected support documentation, depending on maintainer direction.
Written by the indexing model from the issue text.
Description
Summary
ContextCacheConfig is accepted by App.Builder and plumbed all the way to InvocationContext, but nothing in the library ever reads it. Setting it is a silent no-op: no cached content is created, no cachedContent is set on the request, and the full system instruction is sent as tokens on every call.
The context caching docs list Java as supported ("Python v1.15.0+, Java v0.1.0+, Kotlin v0.7.0+"), and RFC #543 was closed with "this feature was successfully implemented in version 0.6.0". I believe that closure was made in error: the commit it points to (12defeed, PR #822) added only the configuration class, and its own message says so — "introduces context caching configuration for apps" (+59 lines in ContextCacheConfig.java, +19/-2 in App.java). PR #823 then added the InvocationContext plumbing. No implementation PR followed.
Because the config is accepted, documented as supported, and plumbed to the point of use, there is no signal that it does nothing — you only find out from a cache-hit rate that never moves.
The chain, traced on main
App.Builder.contextCacheConfig(cfg)stores it onApp.Runner.Builder.build()readsapp.contextCacheConfig()(Runner.java:176) and assigns theRunnerfield.Runner.newInvocationContextBuilder()passes it toInvocationContext.Builder.contextCacheConfig(...).InvocationContext.Builder(InvocationContext)copies it, so child contexts carry it.- Reads of that field:
equals/hashCode, and the public getterInvocationContext.contextCacheConfig()(InvocationContext.java:303). - That getter has no callers. The trail ends there.
Supporting checks on main:
- Zero occurrences of
CreateCachedContentConfig,CachedContent,caches.createorCacheMetadataanywhere in the repo. ContextCacheConfig.getTtlString()— javadoc'd "Returns TTL as string format for cache creation", i.e. written specifically for this feature — has no callers either.- The only cache-adjacent code reads
usageMetadata.cachedContentTokenCountback off responses for reporting (GeminiUtil,BigQueryAgentAnalyticsPlugin), which is a different thing. - Verified identically against the published 1.9.0 and 1.10.1 sources jars, so this is not new.
Reproduction
Configure caching on an App, run one turn against a recording BaseLlm, and inspect the request:
ContextCacheConfig caching =
new ContextCacheConfig(/* maxInvocations= */ 10, Duration.ofHours(1), /* minTokens= */ 0);
App app = App.builder()
.name("repro")
.rootAgent(LlmAgent.builder().name("repro").model(recordingLlm)
.instruction("...32KB of static instruction...").build())
.contextCacheConfig(caching)
.build();
Runner runner = Runner.builder()
.app(app)
.sessionService(new InMemorySessionService())
.artifactService(new InMemoryArtifactService())
.build();
// run one turn, then on the LlmRequest the model received:
assertThat(request.config().get().cachedContent()).isEmpty(); // passes
assertThat(request.config().get().systemInstruction()).isPresent(); // passes — full text, inline
assertThat(app.contextCacheConfig()).isEqualTo(caching); // passes — config was kept
Observed, with minTokens = 0 and a ~32KB instruction so no size floor applies:
| Prompt supplied as | cachedContent on the request |
System instruction |
|---|---|---|
Instruction.Provider |
empty | inline, full text |
static .instruction(String) |
empty | inline, full text |
systemInstruction on generateContentConfig |
empty | inline, full text |
The request is byte-identical with the config present and absent. Response-side too: a response reporting no cachedContentTokenCount reaches the event with none invented for it, and one reporting cachedContentTokenCount passes through untouched (BaseLlmFlow copies usageMetadata verbatim) — so the reporting channel works and the absence above is genuine.
Impact
For an agent with a large static instruction this is the whole cost saving silently not happening. In our case the evaluator prompt is ~7k tokens, about two thirds of every request's input tokens. Implicit caching does not substitute for it at low request rates — entries live a few minutes, so at single-digit requests per minute spread over several prompt variants we measured roughly a 1% hit rate.
Suggested resolution
Either would resolve the confusion; the first is obviously preferable:
- Implement it, porting
gemini_context_cache_manager.py,cache_metadata.pyandflows/llm_flows/context/_cache.pyfrom adk-python. I'm happy to attempt this as a PR if that would be welcome, but given the size (~1k lines of Python, with non-trivial fingerprint/invalidation state) and that this is a design-sensitive area, I'd rather check first whether it's already planned internally or whether there's a preferred shape — the package-structure question from #543 was never settled publicly. - Failing that, make the no-op loud — reopen #543, correct the docs' Java support claim, and either javadoc
ContextCacheConfigas not-yet-implemented on Java or log a warning when anAppis built with one. That costs little and stops the next person spending a day proving the negative.
Happy to supply the full repro as a test case if useful.
Environment
- adk-java
main, plus published 1.9.0 and 1.10.1 (behaviour identical) - Java 25,
google-genai1.58.0
- Dominant language
- Java
- Stars
- 1.7k
- Forks
- 431
- Avg merge
- 3d 2h
- Merged PRs (30d)
- 46
Getting set up
Starts the project's dev container in your browser, under your own GitHub account.
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from google/adk-java
-
GeminiUtil placeholder user turn ("Continue output. DO NOT look at this line ...") is flagged by prompt injection filtersPossibly taken @innoprej claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
[spring-ai] ToolConverter silently drops enum and items from tool parameter schemasPossibly taken @hirematha claimed this 3 days ago. Openneeds review
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
google/adk-java#1609 · 2 comments · 1 assignee ·
Maintainers usually reply within 1 day
-
[spring-ai] Streaming responses ending with CJK punctuation (。!?) are misclassified as partial and never persisted to the sessionPossibly taken @hirematha claimed this 3 days ago. Openwaiting on reporter
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
google/adk-java#1608 · 2 comments · 1 assignee ·
Maintainers usually reply within 1 day
-
[core] Client disconnects don't cancel the model stream (per-step flow is cached) — and there is no public API to cancel an in-flight runPossibly taken @hemasekhar-p claimed this 2 days ago. Openneeds review
google/adk-java#1618 · 6 comments · 1 assignee ·
Maintainers usually reply within 1 day
-
[spring-ai] Bridge drops reasoning_content (thinking) — surface it as partial events and/or persist itPossibly taken @hemasekhar-p claimed this 2 days ago. Openneeds review
google/adk-java#1616 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
Similar issues
-
waiting-for-triage
Difficulty 1/5 Under an hour Newbie friendliness 72/100
spring-cloud/spring-cloud-openfeign#1443 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 1-3 hours Newbie friendliness 84/100
ADORSYS-GIS/keycloak-oid4vp-plugin#221 ·
Maintainers usually reply within 2 days
-
status: team-only type: dependency-upgrade
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
spring-projects/spring-boot#52099 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 67/100
tchiotludo/akhq#3307 · 1 reaction ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
objectionary/jeo-maven-plugin#1885 ·
Maintainers usually reply within 4 days