[Feature Request] Add opt-in retry policy for transient LLM failures
Maintainers usually reply within 1 day
Assessment
This issue has not been assessed yet.
Description
Add opt-in retry policy for transient LLM failures
Please make sure you read the contribution guide and file the issues in the right place.
Contribution guide
🔴 Required Information
Is your feature request related to a specific problem?
Yes.
The Java ADK Runner currently does not provide a configurable retry mechanism for transient LLM provider failures such as rate limiting, service unavailability, network timeouts, or connection failures.
Applications must either fail the invocation immediately or retry the entire Runner.runAsync(...) call. Retrying the complete invocation is unsafe because the runner may have already:
- Appended the user message to the session.
- Updated session state.
- Emitted or persisted model events.
- Executed tools with external side effects.
A whole-invocation retry can therefore duplicate session events or repeat non-idempotent tool operations.
Describe the Solution You'd Like
Add an opt-in retry policy at the LLM model-call boundary inside BaseLlmFlow.
The policy should wrap BaseLlm.generateContent(...), rather than retrying the complete runner invocation. This would allow transient provider failures to be retried without replaying user messages, tools, or other runner-level side effects.
The proposed behavior is:
- Retry is disabled by default to preserve existing behavior.
- Retry transient provider and network failures only.
- Use configurable bounded exponential backoff with optional jitter.
- Count every provider attempt against
RunConfig.maxLlmCalls(). - Never retry after a streaming response has emitted its first
LlmResponse. - Delegate provider-specific exception classification to each
BaseLlmimplementation. - Record retry attempts in logs and tracing attributes.
Impact on your work
We use ADK Java for event-driven agent workflows where temporary rate limits, provider outages, and network interruptions should not immediately fail an otherwise recoverable invocation.
Runner-level retries are not suitable because our agents can update session history and invoke tools with external side effects. Retrying only the model call would improve resilience while preserving session consistency and avoiding duplicate operations.
This feature would also let us enforce a predictable cost limit because every failed and successful provider attempt would count toward maxLlmCalls().
Timeline: N/A.
Willingness to contribute
Yes. We are willing to implement this feature and submit a PR.
🟡 Recommended Information
Describe Alternatives You've Considered
-
Retrying
Runner.runAsync(...)This can duplicate user events, session mutations, model events, and tool executions because the runner appends the user message before agent execution begins.
-
Retrying in an application-specific agent wrapper
This avoids modifying ADK but does not provide a reusable solution for other ADK Java users. It also operates at too broad a boundary and can repeat agent-side effects.
-
Relying entirely on provider SDK retries
Provider behavior varies, is not consistently configurable across supported LLMs, and may not integrate with ADK's
maxLlmCalls()cost limit or tracing. -
Retrying all exceptions
This could repeatedly submit invalid prompts, schemas, credentials, or model configurations. Retry decisions should remain provider-specific and limited to transient failures.
Proposed API / Implementation
Add an opt-in retry configuration to RunConfig:
RunConfig runConfig =
RunConfig.builder()
.maxLlmCalls(12)
.retryConfig(
RetryConfig.builder()
.maxAttempts(3)
.initialBackoff(Duration.ofMillis(500))
.maxBackoff(Duration.ofSeconds(8))
.multiplier(2.0)
.jitterRatio(0.2)
.retryableStatusCodes(Set.of(408, 429, 500, 502, 503, 504))
.build())
.build();
The default configuration would disable retries:
RetryConfig.disabled(); // maxAttempts = 1
Add a provider-specific classification method to BaseLlm:
public abstract boolean isExceptionRetryable(
Throwable exception,
Set<Integer> retryableStatusCodes);
Each implementation would classify its own SDK exceptions:
public class Gemini extends BaseLlm {
@Override
public boolean isExceptionRetryable(
Throwable exception,
Set<Integer> retryableStatusCodes) {
return exception instanceof ApiException apiException
&& retryableStatusCodes.contains(apiException.code())
|| exception instanceof GenAiIOException
|| exception instanceof IOException
|| exception instanceof TimeoutException;
}
}
Equivalent implementations would be provided for OpenAI, Groq, Claude, and Apigee.
BaseLlmFlow would delegate the provider call to a shared retry policy:
return LlmRetryPolicy.execute(
context,
llm,
finalLlmRequest,
context.runConfig().streamingMode() == StreamingMode.SSE);
The shared policy would:
- Increment the LLM call count before each real provider attempt.
- Call
BaseLlm.generateContent(...). - Traverse the exception cause chain safely.
- Ask
llm.isExceptionRetryable(...)to classify each cause. - Apply bounded exponential backoff and jitter.
- Stop after
maxAttempts. - Stop immediately after any streaming response has been emitted.
- Propagate
LlmCallsLimitExceededExceptionwithout retrying.
Example classification flow:
static boolean isRetryable(
BaseLlm llm,
Throwable error,
RetryConfig config) {
Set<Throwable> visited =
Collections.newSetFromMap(new IdentityHashMap<>());
for (Throwable current = error;
current != null && visited.add(current);
current = current.getCause()) {
if (current instanceof LlmCallsLimitExceededException) {
return false;
}
if (llm.isExceptionRetryable(
current, config.retryableStatusCodes())) {
return true;
}
}
return false;
}
Additional Context
The intended retry boundary is:
Runner
→ append user event once
→ agent.runAsync(...)
→ BaseLlmFlow
→ LlmRetryPolicy
→ BaseLlm.generateContent(...)
Failed attempts would not produce ADK events or enter session history.
Suggested acceptance criteria:
- Existing callers experience no behavior change unless retries are enabled.
- Transient provider failures are retried with bounded backoff.
- Non-retryable client, authentication, and configuration errors fail immediately.
- Failed attempts do not create persisted model events.
- User messages and tool calls are not replayed.
- Every provider attempt consumes the
maxLlmCalls()budget. - Streaming calls retry only when no response has been emitted.
- Retry attempts include model, agent, attempt number, delay, and error type in logs or traces.
- Unit tests cover disabled retries, success after retry, non-retryable failures, exhausted attempts, call-budget exhaustion, and partial-stream failures.
- Dominant language
- Java
- Stars
- 1.7k
- Forks
- 433
- Avg merge
- 3d 11h
- Merged PRs (30d)
- 50
Getting set up
Starts the project's dev container in your browser, under your own GitHub account.
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from google/adk-java
-
GeminiUtil placeholder user turn ("Continue output. DO NOT look at this line ...") is flagged by prompt injection filtersPossibly taken @hemasekhar-p claimed this 1 day ago. Openneeds review
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
google/adk-java#1628 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
-
[spring-ai] ToolConverter silently drops enum and items from tool parameter schemasPossibly taken @hirematha claimed this 4 days ago. Openneeds review
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
google/adk-java#1609 · 2 comments · 1 assignee ·
Maintainers usually reply within 1 day
-
[spring-ai] Streaming responses ending with CJK punctuation (。!?) are misclassified as partial and never persisted to the sessionPossibly taken @hirematha claimed this 4 days ago. Openneeds review
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
google/adk-java#1608 · 3 comments · 1 assignee ·
Maintainers usually reply within 1 day
-
Claude model throws UnsupportedOperationException("Not supported yet.") on thinking blocks from Claude 5 modelsPossibly taken @hemasekhar-p claimed this 1 day ago. Openneeds review
google/adk-java#1630 · 2 comments · 1 assignee ·
Maintainers usually reply within 1 day
-
[core] Client disconnects don't cancel the model stream (per-step flow is cached) — and there is no public API to cancel an in-flight runPossibly taken @hemasekhar-p claimed this 3 days ago. Openneeds review
google/adk-java#1618 · 6 comments · 1 assignee ·
Maintainers usually reply within 1 day
Similar issues
-
Fix Math.ceilDiv wrong result for exact positive divisionsPossibly taken @pamod-madubashana claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
scala-native/scala-native#5094 ·
Maintainers usually reply within 1 day
-
[Bug] AI unread message badge counts a batch of new bubbles as one messagePossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
apache/rocketmq-dashboard#5784 ·
Maintainers usually reply within 3 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
PCL-Community/PCL-CE#3658 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 66/100
apache/skywalking#14127 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day