[Bug]: Push notification store failure rewrites a completed task as FAILED (DefaultRequestHandlerV2)
メンテナーはふだん 2 日以内に返信
@rohityan がすでに取り組んでいます。
2026年10月5日 から。
評価
この issue はまだ評価されていません。
説明
What happened?
In DefaultRequestHandlerV2, a failure in the push-notification infrastructure (not in the agent) rewrites an otherwise successful task to TASK_STATE_FAILED and surfaces an error to the client instead of the task result.
In the reproduction below, the agent does everything right — it creates the task and completes it via TaskUpdater (which stamps status.timestamp). The only injected failure is a transient exception from PushNotificationConfigStore.get_info_for_dispatch (e.g. a database hiccup while reading webhook configs). Observed:
--- CONTROL: healthy push config store ---
on_message_send -> Task (state=TASK_STATE_COMPLETED)
persisted task state: TASK_STATE_COMPLETED, timestamp set: True
--- BUG CASE: config store read fails ---
on_message_send RAISED RuntimeError: transient DB error while reading push configs
persisted task state: TASK_STATE_FAILED, timestamp set: False
The task the agent completed is persisted as FAILED, and message/send raises instead of returning the Task.
Root cause
Two layers of missing error boundary:
BasePushNotificationSender.send_notification(src/a2a/server/tasks/base_push_notification_sender.py):get_info_for_dispatch(line 72) and thepush_url_validatorcall (lines 93–98) run outside thetry:that starts at line 99. The dispatch path itself is properly guarded —_dispatch_notificationcatches everything and returnsbool(lines 127–133) — so the design intent is clearly "push failures must not affect the task", but the config-read and validator escapes that boundary.EventConsumer._update_task_state(src/a2a/server/agent_execution/active_task.py:345-350) awaitssend_notificationinline without any guard. It is called from_process_event_inner(line 235); any exception propagates to the consumer's genericexcept Exception(lines 139–153), which calls_write_failed_status()— rewriting the task toTASK_STATE_FAILEDand enqueueing the error to subscribers.
Secondary observation (visible in the repro output): the framework-written FAILED status carries no timestamp, which also nulls last_updated for DatabaseTaskStore rows.
Reproduction
Self-contained script (a2a-sdk @ main e649325, Python 3.10)
import asyncio
import httpx
from a2a.auth.user import UnauthenticatedUser
from a2a.helpers.proto_helpers import new_task_from_user_message
from a2a.server.agent_execution import AgentExecutor, RequestContext
from a2a.server.context import ServerCallContext
from a2a.server.events import EventQueue
from a2a.server.request_handlers import DefaultRequestHandlerV2
from a2a.server.tasks import (
InMemoryPushNotificationConfigStore,
InMemoryTaskStore,
TaskUpdater,
)
from a2a.server.tasks.base_push_notification_sender import (
BasePushNotificationSender,
)
from a2a.types.a2a_pb2 import (
AgentCapabilities,
AgentCard,
Message,
Part,
Role,
SendMessageConfiguration,
SendMessageRequest,
TaskState,
)
class CompletingExecutor(AgentExecutor):
# The agent does everything right: creates the task, then completes it
# via TaskUpdater (which stamps status.timestamp).
task_id = None
async def execute(self, context, event_queue):
task = new_task_from_user_message(context.message)
self.task_id = task.id
await event_queue.enqueue_event(task)
updater = TaskUpdater(event_queue, task.id, task.context_id)
await updater.update_status(TaskState.TASK_STATE_COMPLETED)
async def cancel(self, context, event_queue):
pass
class FlakyConfigStore(InMemoryPushNotificationConfigStore):
# Simulates a transient infrastructure failure on the dispatch read path.
async def get_info_for_dispatch(self, task_id):
raise RuntimeError('transient DB error while reading push configs')
def build_handler(store):
executor = CompletingExecutor()
handler = DefaultRequestHandlerV2(
agent_executor=executor,
task_store=InMemoryTaskStore(),
push_config_store=store,
push_sender=BasePushNotificationSender(
httpx_client=httpx.AsyncClient(), config_store=store
),
agent_card=AgentCard(
name='t',
version='1.0',
capabilities=AgentCapabilities(
streaming=True, push_notifications=True
),
),
)
return handler, executor
async def run_case(label, store):
handler, executor = build_handler(store)
params = SendMessageRequest(
message=Message(
role=Role.ROLE_USER, message_id='m1', parts=[Part(text='Hi')]
),
configuration=SendMessageConfiguration(accepted_output_modes=['text/plain']),
)
print(f'--- {label} ---')
try:
result = await handler.on_message_send(
params, ServerCallContext(user=UnauthenticatedUser())
)
state = getattr(getattr(result, 'status', None), 'state', None)
print(
f' on_message_send -> {type(result).__name__} '
f'(state={TaskState.Name(state) if state is not None else None})'
)
except Exception as e:
print(f' on_message_send RAISED {type(e).__name__}: {e}')
task = await handler.task_store.get(
executor.task_id, ServerCallContext(user=UnauthenticatedUser())
)
print(
f' persisted task state: {TaskState.Name(task.status.state)}, '
f'timestamp set: {task.status.HasField("timestamp")}'
)
async def main():
await run_case(
'CONTROL: healthy push config store',
InMemoryPushNotificationConfigStore(),
)
await run_case('BUG CASE: config store read fails', FlakyConfigStore())
asyncio.run(main())
Expected behavior
A failure to read push configs or to run the URL validator should degrade to "notification not sent" (logged), and must never mutate the task's state or fail the client request. This matches the intent already expressed by _dispatch_notification's error boundary.
Suggested fix
Two small, non-breaking changes:
- In
EventConsumer._update_task_state, wrap thesend_notificationcall intry/except Exceptionwithlogger.exception(...)— push delivery is best-effort. - In
BasePushNotificationSender.send_notification, move theget_info_for_dispatchcall and the validator invocation inside the error boundary (log and return early on failure).
Happy to submit a PR with tests if the approach sounds right.
Relevant log output
RuntimeError: transient DB error while reading push configs
Dispatcher task was cancelled or finished. Events may be lost.
Code of Conduct
- I agree to follow this project's Code of Conduct
- 主要言語
- Python
- スター
- 2.2k
- フォーク
- 509
- 平均マージ
- 3日 11時間
- マージ済み PR(30日)
- 43
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
a2aproject/a2a-python のほかの issue
-
[Bug]: REST task/request id sanitization対応中かも @Linux2010 が 101 日前に担当しました。 オープンmaintainers-only
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
a2aproject/a2a-python#805 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
[Bug]: After a streamed task is cancelled, its background producer never finishes対応中かも @rohityan が今日担当しました。 オープン
a2aproject/a2a-python#1322 · 担当者 2 名 ·
メンテナーはふだん 2 日以内に返信
-
v0.3 gRPC and REST SendMessage without configuration run non-blocking対応中かも @rohityan が 1 日前に担当しました。 オープン
a2aproject/a2a-python#1321 · 担当者 2 名 ·
メンテナーはふだん 2 日以内に返信
-
v0.3 JSON-RPC and REST GetTask return no history when history_length is 0対応中かも @rohityan が 3 日前に担当しました。 オープン
a2aproject/a2a-python#1311 · 担当者 2 名 ·
メンテナーはふだん 2 日以内に返信
-
[Bug]: DefaultRequestHandlerV2 keeps the ActiveTask (producer, consumer, 2 dispatchers) alive forever after a direct Message or input-required response対応中かも @rohityan が 5 日前に担当しました。 オープン
a2aproject/a2a-python#1296 · 担当者 2 名 ·
メンテナーはふだん 2 日以内に返信
a2aproject/a2a-python の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
-
Task
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
war-and-code/dircue#200 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 87/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 84/100
メンテナーはふだん 1 日以内に返信