[Bug]: Push notification store failure rewrites a completed task as FAILED (DefaultRequestHandlerV2)
维护者通常 2 天内回复
@rohityan 已经在做这个了。
开始于 2026年10月5日。
评估
这个 Issue 还没有评估数据。
描述
What happened?
In DefaultRequestHandlerV2, a failure in the push-notification infrastructure (not in the agent) rewrites an otherwise successful task to TASK_STATE_FAILED and surfaces an error to the client instead of the task result.
In the reproduction below, the agent does everything right — it creates the task and completes it via TaskUpdater (which stamps status.timestamp). The only injected failure is a transient exception from PushNotificationConfigStore.get_info_for_dispatch (e.g. a database hiccup while reading webhook configs). Observed:
--- CONTROL: healthy push config store ---
on_message_send -> Task (state=TASK_STATE_COMPLETED)
persisted task state: TASK_STATE_COMPLETED, timestamp set: True
--- BUG CASE: config store read fails ---
on_message_send RAISED RuntimeError: transient DB error while reading push configs
persisted task state: TASK_STATE_FAILED, timestamp set: False
The task the agent completed is persisted as FAILED, and message/send raises instead of returning the Task.
Root cause
Two layers of missing error boundary:
BasePushNotificationSender.send_notification(src/a2a/server/tasks/base_push_notification_sender.py):get_info_for_dispatch(line 72) and thepush_url_validatorcall (lines 93–98) run outside thetry:that starts at line 99. The dispatch path itself is properly guarded —_dispatch_notificationcatches everything and returnsbool(lines 127–133) — so the design intent is clearly "push failures must not affect the task", but the config-read and validator escapes that boundary.EventConsumer._update_task_state(src/a2a/server/agent_execution/active_task.py:345-350) awaitssend_notificationinline without any guard. It is called from_process_event_inner(line 235); any exception propagates to the consumer's genericexcept Exception(lines 139–153), which calls_write_failed_status()— rewriting the task toTASK_STATE_FAILEDand enqueueing the error to subscribers.
Secondary observation (visible in the repro output): the framework-written FAILED status carries no timestamp, which also nulls last_updated for DatabaseTaskStore rows.
Reproduction
Self-contained script (a2a-sdk @ main e649325, Python 3.10)
import asyncio
import httpx
from a2a.auth.user import UnauthenticatedUser
from a2a.helpers.proto_helpers import new_task_from_user_message
from a2a.server.agent_execution import AgentExecutor, RequestContext
from a2a.server.context import ServerCallContext
from a2a.server.events import EventQueue
from a2a.server.request_handlers import DefaultRequestHandlerV2
from a2a.server.tasks import (
InMemoryPushNotificationConfigStore,
InMemoryTaskStore,
TaskUpdater,
)
from a2a.server.tasks.base_push_notification_sender import (
BasePushNotificationSender,
)
from a2a.types.a2a_pb2 import (
AgentCapabilities,
AgentCard,
Message,
Part,
Role,
SendMessageConfiguration,
SendMessageRequest,
TaskState,
)
class CompletingExecutor(AgentExecutor):
# The agent does everything right: creates the task, then completes it
# via TaskUpdater (which stamps status.timestamp).
task_id = None
async def execute(self, context, event_queue):
task = new_task_from_user_message(context.message)
self.task_id = task.id
await event_queue.enqueue_event(task)
updater = TaskUpdater(event_queue, task.id, task.context_id)
await updater.update_status(TaskState.TASK_STATE_COMPLETED)
async def cancel(self, context, event_queue):
pass
class FlakyConfigStore(InMemoryPushNotificationConfigStore):
# Simulates a transient infrastructure failure on the dispatch read path.
async def get_info_for_dispatch(self, task_id):
raise RuntimeError('transient DB error while reading push configs')
def build_handler(store):
executor = CompletingExecutor()
handler = DefaultRequestHandlerV2(
agent_executor=executor,
task_store=InMemoryTaskStore(),
push_config_store=store,
push_sender=BasePushNotificationSender(
httpx_client=httpx.AsyncClient(), config_store=store
),
agent_card=AgentCard(
name='t',
version='1.0',
capabilities=AgentCapabilities(
streaming=True, push_notifications=True
),
),
)
return handler, executor
async def run_case(label, store):
handler, executor = build_handler(store)
params = SendMessageRequest(
message=Message(
role=Role.ROLE_USER, message_id='m1', parts=[Part(text='Hi')]
),
configuration=SendMessageConfiguration(accepted_output_modes=['text/plain']),
)
print(f'--- {label} ---')
try:
result = await handler.on_message_send(
params, ServerCallContext(user=UnauthenticatedUser())
)
state = getattr(getattr(result, 'status', None), 'state', None)
print(
f' on_message_send -> {type(result).__name__} '
f'(state={TaskState.Name(state) if state is not None else None})'
)
except Exception as e:
print(f' on_message_send RAISED {type(e).__name__}: {e}')
task = await handler.task_store.get(
executor.task_id, ServerCallContext(user=UnauthenticatedUser())
)
print(
f' persisted task state: {TaskState.Name(task.status.state)}, '
f'timestamp set: {task.status.HasField("timestamp")}'
)
async def main():
await run_case(
'CONTROL: healthy push config store',
InMemoryPushNotificationConfigStore(),
)
await run_case('BUG CASE: config store read fails', FlakyConfigStore())
asyncio.run(main())
Expected behavior
A failure to read push configs or to run the URL validator should degrade to "notification not sent" (logged), and must never mutate the task's state or fail the client request. This matches the intent already expressed by _dispatch_notification's error boundary.
Suggested fix
Two small, non-breaking changes:
- In
EventConsumer._update_task_state, wrap thesend_notificationcall intry/except Exceptionwithlogger.exception(...)— push delivery is best-effort. - In
BasePushNotificationSender.send_notification, move theget_info_for_dispatchcall and the validator invocation inside the error boundary (log and return early on failure).
Happy to submit a PR with tests if the approach sounds right.
Relevant log output
RuntimeError: transient DB error while reading push configs
Dispatcher task was cancelled or finished. Events may be lost.
Code of Conduct
- I agree to follow this project's Code of Conduct
- 主要语言
- Python
- 星标
- 2.2k
- 派生
- 509
- 平均合并
- 3 天 18 小时
- 30 天内合并 PR
- 44
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
a2aproject/a2a-python 的其他 Issue
-
[Bug]: REST task/request id sanitization可能已有人在做 @Linux2010 于 103 天前认领。 未关闭maintainers-only
难度 2/5 1-3 小时 新手友好度 65/100
a2aproject/a2a-python#805 · 1 条评论 ·
维护者通常 2 天内回复
-
[Feat]: Cluster mode: detect and recover tasks abandoned by a crashed replica (heartbeat/lease)可能已有人在做 @rohityan 于 2 天前认领。 未关闭
a2aproject/a2a-python#1324 · 已指派 1 人 ·
维护者通常 2 天内回复
-
[Bug]: Cluster mode: SubscribeToTask on a replica holding a paused ActiveTask returns a stale INPUT_REQUIRED snapshot and closes可能已有人在做 @rohityan 于 2 天前认领。 未关闭component: server
a2aproject/a2a-python#1323 · 已指派 1 人 ·
维护者通常 2 天内回复
-
[Bug]: After a streamed task is cancelled, its background producer never finishes可能已有人在做 @rohityan 于 2 天前认领。 未关闭component: server status:awaiting response
a2aproject/a2a-python#1322 · 1 条评论 · 已指派 1 人 ·
维护者通常 2 天内回复
-
v0.3 gRPC and REST SendMessage without configuration run non-blocking可能已有人在做 @rohityan 于 3 天前认领。 未关闭question status:awaiting response
a2aproject/a2a-python#1321 · 1 条评论 · 已指派 1 人 ·
维护者通常 2 天内回复
查看 a2aproject/a2a-python 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 82/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 62/100
-
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
[BUG] Multi-day events show "Ended" while still in progress可能已有人在做 @tarunagnihotri534 今天认领。 未关闭bug
难度 2/5 1-3 小时 新手友好度 85/100
data-umbrella/du-event-board#231 · 2 条评论 ·
-
avl_automation: the generated control surface block isn't valid XML (typo in avl_out_parse.py)可能已有人在做 @brksol 今天认领。 未关闭
难度 1/5 1 小时以内 新手友好度 92/100
PX4/PX4-gazebo-models#164 ·