[Feature Request] Add opt-in retry policy for transient LLM failures
维护者通常 1 天内回复
评估
这个 Issue 还没有评估数据。
描述
Add opt-in retry policy for transient LLM failures
Please make sure you read the contribution guide and file the issues in the right place.
Contribution guide
🔴 Required Information
Is your feature request related to a specific problem?
Yes.
The Java ADK Runner currently does not provide a configurable retry mechanism for transient LLM provider failures such as rate limiting, service unavailability, network timeouts, or connection failures.
Applications must either fail the invocation immediately or retry the entire Runner.runAsync(...) call. Retrying the complete invocation is unsafe because the runner may have already:
- Appended the user message to the session.
- Updated session state.
- Emitted or persisted model events.
- Executed tools with external side effects.
A whole-invocation retry can therefore duplicate session events or repeat non-idempotent tool operations.
Describe the Solution You'd Like
Add an opt-in retry policy at the LLM model-call boundary inside BaseLlmFlow.
The policy should wrap BaseLlm.generateContent(...), rather than retrying the complete runner invocation. This would allow transient provider failures to be retried without replaying user messages, tools, or other runner-level side effects.
The proposed behavior is:
- Retry is disabled by default to preserve existing behavior.
- Retry transient provider and network failures only.
- Use configurable bounded exponential backoff with optional jitter.
- Count every provider attempt against
RunConfig.maxLlmCalls(). - Never retry after a streaming response has emitted its first
LlmResponse. - Delegate provider-specific exception classification to each
BaseLlmimplementation. - Record retry attempts in logs and tracing attributes.
Impact on your work
We use ADK Java for event-driven agent workflows where temporary rate limits, provider outages, and network interruptions should not immediately fail an otherwise recoverable invocation.
Runner-level retries are not suitable because our agents can update session history and invoke tools with external side effects. Retrying only the model call would improve resilience while preserving session consistency and avoiding duplicate operations.
This feature would also let us enforce a predictable cost limit because every failed and successful provider attempt would count toward maxLlmCalls().
Timeline: N/A.
Willingness to contribute
Yes. We are willing to implement this feature and submit a PR.
🟡 Recommended Information
Describe Alternatives You've Considered
-
Retrying
Runner.runAsync(...)This can duplicate user events, session mutations, model events, and tool executions because the runner appends the user message before agent execution begins.
-
Retrying in an application-specific agent wrapper
This avoids modifying ADK but does not provide a reusable solution for other ADK Java users. It also operates at too broad a boundary and can repeat agent-side effects.
-
Relying entirely on provider SDK retries
Provider behavior varies, is not consistently configurable across supported LLMs, and may not integrate with ADK's
maxLlmCalls()cost limit or tracing. -
Retrying all exceptions
This could repeatedly submit invalid prompts, schemas, credentials, or model configurations. Retry decisions should remain provider-specific and limited to transient failures.
Proposed API / Implementation
Add an opt-in retry configuration to RunConfig:
RunConfig runConfig =
RunConfig.builder()
.maxLlmCalls(12)
.retryConfig(
RetryConfig.builder()
.maxAttempts(3)
.initialBackoff(Duration.ofMillis(500))
.maxBackoff(Duration.ofSeconds(8))
.multiplier(2.0)
.jitterRatio(0.2)
.retryableStatusCodes(Set.of(408, 429, 500, 502, 503, 504))
.build())
.build();
The default configuration would disable retries:
RetryConfig.disabled(); // maxAttempts = 1
Add a provider-specific classification method to BaseLlm:
public abstract boolean isExceptionRetryable(
Throwable exception,
Set<Integer> retryableStatusCodes);
Each implementation would classify its own SDK exceptions:
public class Gemini extends BaseLlm {
@Override
public boolean isExceptionRetryable(
Throwable exception,
Set<Integer> retryableStatusCodes) {
return exception instanceof ApiException apiException
&& retryableStatusCodes.contains(apiException.code())
|| exception instanceof GenAiIOException
|| exception instanceof IOException
|| exception instanceof TimeoutException;
}
}
Equivalent implementations would be provided for OpenAI, Groq, Claude, and Apigee.
BaseLlmFlow would delegate the provider call to a shared retry policy:
return LlmRetryPolicy.execute(
context,
llm,
finalLlmRequest,
context.runConfig().streamingMode() == StreamingMode.SSE);
The shared policy would:
- Increment the LLM call count before each real provider attempt.
- Call
BaseLlm.generateContent(...). - Traverse the exception cause chain safely.
- Ask
llm.isExceptionRetryable(...)to classify each cause. - Apply bounded exponential backoff and jitter.
- Stop after
maxAttempts. - Stop immediately after any streaming response has been emitted.
- Propagate
LlmCallsLimitExceededExceptionwithout retrying.
Example classification flow:
static boolean isRetryable(
BaseLlm llm,
Throwable error,
RetryConfig config) {
Set<Throwable> visited =
Collections.newSetFromMap(new IdentityHashMap<>());
for (Throwable current = error;
current != null && visited.add(current);
current = current.getCause()) {
if (current instanceof LlmCallsLimitExceededException) {
return false;
}
if (llm.isExceptionRetryable(
current, config.retryableStatusCodes())) {
return true;
}
}
return false;
}
Additional Context
The intended retry boundary is:
Runner
→ append user event once
→ agent.runAsync(...)
→ BaseLlmFlow
→ LlmRetryPolicy
→ BaseLlm.generateContent(...)
Failed attempts would not produce ADK events or enter session history.
Suggested acceptance criteria:
- Existing callers experience no behavior change unless retries are enabled.
- Transient provider failures are retried with bounded backoff.
- Non-retryable client, authentication, and configuration errors fail immediately.
- Failed attempts do not create persisted model events.
- User messages and tool calls are not replayed.
- Every provider attempt consumes the
maxLlmCalls()budget. - Streaming calls retry only when no response has been emitted.
- Retry attempts include model, agent, attempt number, delay, and error type in logs or traces.
- Unit tests cover disabled retries, success after retry, non-retryable failures, exhausted attempts, call-budget exhaustion, and partial-stream failures.
- 主要语言
- Java
- 星标
- 1.7k
- 派生
- 431
- 平均合并
- 3 天 2 小时
- 30 天内合并 PR
- 46
环境准备
在浏览器里用你自己的 GitHub 账号启动这个项目的开发容器。
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
google/adk-java 的其他 Issue
-
GeminiUtil placeholder user turn ("Continue output. DO NOT look at this line ...") is flagged by prompt injection filters可能已有人在做 @innoprej 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复
-
[spring-ai] ToolConverter silently drops enum and items from tool parameter schemas可能已有人在做 @hirematha 于 3 天前认领。 未关闭needs review
难度 2/5 1-3 小时 新手友好度 76/100
google/adk-java#1609 · 2 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[spring-ai] Streaming responses ending with CJK punctuation (。!?) are misclassified as partial and never persisted to the session可能已有人在做 @hirematha 于 3 天前认领。 未关闭waiting on reporter
难度 2/5 1-3 小时 新手友好度 84/100
google/adk-java#1608 · 2 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[core] Client disconnects don't cancel the model stream (per-step flow is cached) — and there is no public API to cancel an in-flight run可能已有人在做 @hemasekhar-p 于 2 天前认领。 未关闭needs review
google/adk-java#1618 · 6 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[spring-ai] Bridge drops reasoning_content (thinking) — surface it as partial events and/or persist it可能已有人在做 @hemasekhar-p 于 3 天前认领。 未关闭needs review
google/adk-java#1616 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 64/100
-
C21 publishes `reactivemongo/core/SSL` as Java 23 bytecode — TLS connections fail on any JDK < 23未关闭
难度 2/5 1-3 小时 新手友好度 74/100
ReactiveMongo/ReactiveMongo#1520 ·
维护者通常 1 天内回复
-
enhancement
难度 2/5 1-3 小时 新手友好度 65/100
liquid-java/liquidjava#373 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 70/100
ga4gh/phenopacket-schema#465 ·
-
难度 2/5 1-3 小时 新手友好度 72/100
NationalSecurityAgency/ghidra#9748 ·
维护者通常 1 天内回复