[core] Client disconnects don't cancel the model stream (per-step flow is cached) — and there is no public API to cancel an in-flight run
Maintainer thường phản hồi trong vòng 1 ngày
Đánh giá
Issue này chưa được đánh giá.
Mô tả
Describe the Bug
When a client disconnects while an agent is streaming (e.g. the SSE consumer cancels its subscription to Runner's event flow), the underlying model call keeps running to completion. Tokens continue to be billed with no consumer attached. In one production incident the model kept generating for 3+ minutes after the client went away (observed via gateway-side logs and token metering).
The cause appears to be BaseLlmFlow.java line 530 (main at the time of writing):
Flowable<Event> currentStepEvents = runOneStep(spanContext, invocationContext).cache();
With RxJava cache(), the upstream connection is only disposed when all subscribers of the cached flow cancel. As long as any subscriber attached by the framework itself (event handling/aggregation within the invocation) is alive, a user-side cancellation cannot reach the model stream.
On top of that, there is no public API to cancel an in-flight run: neither Runner nor InvocationContext nor the session exposes a "stop generating" handle. The only way to stop a long generation is to let it finish.
To Reproduce
- Serve an agent behind SSE; start a streaming run with a slow/long model response (or a multi-round tool loop).
- Disconnect the SSE client mid-generation and dispose the event-flow subscription.
- Observe on the model/gateway side (or via token metering) that generation continues until the model finishes naturally.
Expected behavior
Cancelling the downstream subscription (or any public cancellation entry point) must reach the model stream: the in-flight LLM request is disposed (or aborted at the next chunk/round boundary), and no further rounds are started.
Environment
- google/adk-java main (1.11.1-SNAPSHOT at the time of writing)
- Spring AI bridge + SSE transport, but the
cache()behavior is in core (BaseLlmFlow) and should affect any transport
Additional context
In production we could only implement "stop generating" with a cooperative cancellation protocol outside the framework: a per-run cancel flag (REST endpoint sets it) that the model-stream loop checks at every chunk and at tool-round boundaries, plus setCancellable hooks for the current round. It works, but it is entirely application-side — the framework neither propagates the cancellation nor offers a canonical way to express it.
Suggested directions (happy to discuss or contribute):
- Replace the unbounded
cache()with cancellation-propagating semantics (e.g.publish().refCount()-style sharing, or a scoped subscription that disposes the source when the invocation ends/cancels), or - Expose a first-class cancellation handle per run/invocation, and document that model loops should check it at chunk/round boundaries (cooperative cancellation), or
- At minimum, document the current behavior prominently — today it silently leaks billable model calls.
- Ngôn ngữ chính
- Java
- Star
- 1.7k
- Fork
- 431
- Merge trung bình
- 3 ngày 2 giờ
- Pull request đã merge (30 ngày)
- 46
Chuẩn bị môi trường
Khởi chạy dev container của dự án ngay trên trình duyệt, bằng tài khoản GitHub của bạn.
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của google/adk-java
-
GeminiUtil placeholder user turn ("Continue output. DO NOT look at this line ...") is flagged by prompt injection filtersCó thể đã có người làm @innoprej đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày
-
[spring-ai] ToolConverter silently drops enum and items from tool parameter schemasCó thể đã có người làm @hirematha đã nhận 3 ngày trước. Đang mởneeds review
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
google/adk-java#1609 · 2 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[spring-ai] Streaming responses ending with CJK punctuation (。!?) are misclassified as partial and never persisted to the sessionCó thể đã có người làm @hirematha đã nhận 3 ngày trước. Đang mởwaiting on reporter
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
google/adk-java#1608 · 2 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[spring-ai] Bridge drops reasoning_content (thinking) — surface it as partial events and/or persist itCó thể đã có người làm @hemasekhar-p đã nhận 2 ngày trước. Đang mởneeds review
google/adk-java#1616 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[FEATURE] Port bypass_multi_tools_limit for built-in search tools from adk-pythonCó thể đã có người làm @hirematha đã nhận 3 ngày trước. Đang mởneeds review
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
google/adk-java#1598 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của google/adk-java
Issue tương tự
-
waiting-for-triage
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 72/100
spring-cloud/spring-cloud-openfeign#1443 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 84/100
ADORSYS-GIS/keycloak-oid4vp-plugin#221 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Upgrade to Spring Pulsar 2.0.8Đang mởstatus: team-only type: dependency-upgrade
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
spring-projects/spring-boot#52099 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 67/100
tchiotludo/akhq#3307 · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
objectionary/jeo-maven-plugin#1885 ·
Maintainer thường phản hồi trong vòng 4 ngày