[Feat]: Cluster mode: detect and recover tasks abandoned by a crashed replica (heartbeat/lease)
Maintainer thường phản hồi trong vòng 2 ngày
@rohityan đang làm issue này rồi.
Từ ngày 7/10/2026.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
Is your feature request related to a problem? Please describe.
Cluster mode (#1281) handles replicas that are alive and disagree: version checks on saves, cancel from any replica, resubscribe from any replica. It does nothing for a replica that dies while running execute() (OOM kill, node loss, a rolling deploy without enough grace time).
When that happens, the task stays non-terminal forever:
GetTaskreturnsTASK_STATE_WORKINGindefinitely.- Every
SubscribeToTaskon another replica keeps pollingtask_eventsand never ends, because no terminal event will ever be written. - Nothing ever re-runs or fails the task. The only way out is a client-issued
CancelTask. - A push-only client never hears anything again.
The client can't tell a dead replica from a slow agent. status.timestamp only moves when the agent emits an event, so a long LLM or tool call looks exactly like a crash.
Detecting abandoned tasks was a stated goal of the Go SDK's distributed mode (a2a-go discussion #139: "Supporting abandoned task detection (e.g. on process reboot)"), and a2a-go ships it (a2a-go #115, a2asrv/workqueue). The Python design discussion (#1224) doesn't cover it.
Describe the solution you'd like
An optional extension point for cluster mode, similar to a2a-go's workqueue.Heartbeater / LeaseManager, along these lines:
- Lease while executing. While
execute()runs, the replica renews a lease (task_id,owner,expires_at, plus the task's owner scope) on an interval. - Release when the turn ends. Release the lease when the task moves into a terminal or interrupted state, so a task waiting on
INPUT_REQUIRED/AUTH_REQUIREDis never treated as abandoned. - Sweep. On every replica, a sweeper atomically claims leases that have expired (for example
FOR UPDATE SKIP LOCKED, so no leader is needed) and applies a policy:fail(default): writeTASK_STATE_FAILED;re-execute: run the turn again, with an attempt counter.
- Write the failure through
VersionedTaskStore.save(..., event=...). This matters in cluster mode:- the event row ends every remote
SubscribeToTaskstream withFAILED; - the version bump fences a replica that wasn't really dead (a long GC pause or a network partition). Its next save gets
ConcurrentTaskModificationError, it re-reads, seesFAILED, and cancels its producer, the same way a remote cancel works today; - push senders can then notify
FAILED.
- the event row ends every remote
Pitfalls we hit running a heartbeat plus sweeper outside the SDK. They're probably worth designing in from the start:
- Release after the durable save, not when
execute()returns. The terminal or interrupted status event may still be on the event queue whenexecute()returns. A crash in that gap leaves a non-terminal task with no lease, which no sweep can ever find. - Release on a state change, not on whatever state is being saved. On a resumed turn, the first save is the held user-message flush, and it re-saves the stale
INPUT_REQUIREDstate (update_with_messagedoesn't changestatus.state). Releasing on "saved an interrupted state" drops the lease at the start of the resumed turn. - A renewal must not recreate a released lease. If renewal is an upsert, a heartbeat that lands just after the release brings the lease back. The sweeper later claims it and fails a task that was correctly paused.
- The sweeper needs the owner scope.
DatabaseTaskStore/VersionedDatabaseTaskStorefilter reads byowner_resolver(context). A sweeper callingget()with an emptyServerCallContextgetsNonefor authenticated tasks and silently does nothing. Store the owner scope with the lease and rebuild the context from it.
Describe alternatives you've considered
- Handle it in the application (what we do today): heartbeat rows, an orphan sweeper, and a
TaskStorewrapper that releases the lease after the durable save. It works, but every deployment has to rebuild it. In cluster mode it also needs the versioned save path and the event table to end remote streams correctly, and those are SDK internals an application shouldn't have to coordinate with. - Leave it out of scope, as a2a-rs did (a2a-rs#149). That's reasonable for single-process servers. But cluster mode is aimed specifically at multi-replica deployments, and replicas dying is the normal failure in that setting.
- A client-side overall timeout is a useful backstop, but the client can't tell which tasks are actually dead.
- The spec-level "Expectations" proposal (A2A#2209) is about business deadlines, not process liveness.
Additional context
- Defaults that worked for us: heartbeat every 30s, a task counts as abandoned after 120s without one, and the sweep runs every 30s. That's roughly 2–2.5 minutes to detect a dead replica.
- Happy to share more detail or help with an implementation if the maintainers think this belongs in the SDK.
Code of Conduct
- I agree to follow this project's Code of Conduct
- Ngôn ngữ chính
- Python
- Star
- 2.2k
- Fork
- 509
- Merge trung bình
- 3 ngày 18 giờ
- Pull request đã merge (30 ngày)
- 44
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của a2aproject/a2a-python
-
[Bug]: REST task/request id sanitizationCó thể đã có người làm @Linux2010 đã nhận 103 ngày trước. Đang mởmaintainers-only
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
a2aproject/a2a-python#805 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug]: Cluster mode: SubscribeToTask on a replica holding a paused ActiveTask returns a stale INPUT_REQUIRED snapshot and closesCó thể đã có người làm @rohityan đã nhận 2 ngày trước. Đang mởcomponent: server
a2aproject/a2a-python#1323 · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug]: After a streamed task is cancelled, its background producer never finishesCó thể đã có người làm @rohityan đã nhận 2 ngày trước. Đang mởcomponent: server status:awaiting response
a2aproject/a2a-python#1322 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
v0.3 gRPC and REST SendMessage without configuration run non-blockingCó thể đã có người làm @rohityan đã nhận 3 ngày trước. Đang mởquestion status:awaiting response
a2aproject/a2a-python#1321 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug]: Push notification store failure rewrites a completed task as FAILED (DefaultRequestHandlerV2)Có thể đã có người làm @rohityan đã nhận 4 ngày trước. Đang mởcomponent: server status:awaiting response
a2aproject/a2a-python#1313 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của a2aproject/a2a-python
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 85/100
MystenLabs/MemWal#1163 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
infertopics leaves new nodes without a topic when untopiced neighbours outnumber topiced onesCó thể đã có người làm @moneebullah25 đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
FinanceFlash/unvibecode#218 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
NVIDIA/earth2studio#1241 ·
Maintainer thường phản hồi trong vòng 3 ngày