[Bug]: QueueShutDown spans still end as ERROR (and exceptions are recorded twice) because start_as_current_span sets status on exception
メンテナーはふだん 2 日以内に返信
評価
この issue はまだ評価されていません。
説明
What happened?
EventQueueSource.dequeue_event spans still end with status ERROR (AsyncQueueShutDown: ) on every normal queue teardown, even though #1075 (fix for #1065) added QueueShutDown to _NON_ERROR_EXCEPTIONS.
trace_function opens its span with OTel's defaults:
with tracer.start_as_current_span(actual_span_name, kind=kind) as span:
start_as_current_span defaults to record_exception=True and set_status_on_exception=True. So when the wrapper re-raises, OTel's use_span records the exception and calls span.set_status(ERROR, ...) for any Exception that leaves the block (opentelemetry/trace/__init__.py, use_span).
QueueShutDownstill ends as ERROR. Theexcept _NON_ERROR_EXCEPTIONSarm correctly leaves the status alone and re-raises. OTel then marks the span ERROR anyway.- Every
Exceptionis recorded twice. Bothexceptarms callspan.record_exception, and OTel records it a second time. So each such span carries twoexceptionevents.
CancelledError isn't affected, because it derives from BaseException and OTel only catches Exception.
The existing test (test_trace_function_async_non_error_exception_does_not_mark_span_error) uses a mocked tracer, so it only sees the wrapper's own set_status calls and not what OTel does on exit.
Expected: a span that ends with QueueShutDown is not ERROR, and a failing span has a single exception event.
Reproduction (a2a-sdk 1.2.1, opentelemetry-sdk 1.42.1, Python 3.12; same code on main at 6ff0a82):
import asyncio
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
from a2a.utils._async_queue_compat import QueueShutDown
from a2a.utils.telemetry import trace_function
@trace_function(span_name="shutdown")
async def shutdown():
raise QueueShutDown()
@trace_function(span_name="boom")
async def boom():
raise ValueError("real failure")
async def main():
for f in (shutdown, boom):
try:
await f()
except Exception:
pass
asyncio.run(main())
for s in exporter.get_finished_spans():
print(s.name, s.status.status_code.name, repr(s.status.description), [e.name for e in s.events])
Output:
shutdown ERROR 'AsyncQueueShutDown: ' ['exception', 'exception']
boom ERROR 'ValueError: real failure' ['exception', 'exception']
Both spans carry the exception twice, and boom's description is OTel's (ValueError: real failure), which overwrote the wrapper's str(e).
With the proposed fix, the same script prints:
shutdown UNSET None ['exception']
boom ERROR 'real failure' ['exception']
We noticed it on an A2A server running DefaultRequestHandlerV2: every streaming request leaves one ERROR dequeue_event span, which shows up as noise in the error rate of the trace backend.
Proposed fix: pass record_exception=False, set_status_on_exception=False to start_as_current_span in both wrappers, since the wrapper already does both itself. PR to follow.
Relevant log output
"name": "a2a.server.events.event_queue_v2.EventQueueSource.dequeue_event",
"kind": "SpanKind.SERVER",
"status": {
"status_code": "ERROR",
"description": "QueueShutDown: "
},
"events": [
{
"name": "exception",
"attributes": {
"exception.type": "culsans.QueueShutDown",
Code of Conduct
- I agree to follow this project's Code of Conduct
- 主要言語
- Python
- スター
- 2.2k
- フォーク
- 509
- 平均マージ
- 3日 11時間
- マージ済み PR(30日)
- 43
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
a2aproject/a2a-python のほかの issue
-
[Bug]: REST task/request id sanitization対応中かも @Linux2010 が 101 日前に担当しました。 オープンmaintainers-only
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
a2aproject/a2a-python#805 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
[Bug]: After a streamed task is cancelled, its background producer never finishes対応中かも @rohityan が今日担当しました。 オープン
a2aproject/a2a-python#1322 · 担当者 2 名 ·
メンテナーはふだん 2 日以内に返信
-
v0.3 gRPC and REST SendMessage without configuration run non-blocking対応中かも @rohityan が今日担当しました。 オープン
a2aproject/a2a-python#1321 · 担当者 2 名 ·
メンテナーはふだん 2 日以内に返信
-
[Bug]: Push notification store failure rewrites a completed task as FAILED (DefaultRequestHandlerV2)対応中かも @rohityan が 2 日前に担当しました。 オープン
a2aproject/a2a-python#1313 · 担当者 2 名 ·
メンテナーはふだん 2 日以内に返信
-
v0.3 JSON-RPC and REST GetTask return no history when history_length is 0対応中かも @rohityan が 2 日前に担当しました。 オープン
a2aproject/a2a-python#1311 · 担当者 2 名 ·
メンテナーはふだん 2 日以内に返信
a2aproject/a2a-python の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 83/100
メンテナーはふだん 1 日以内に返信
-
enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
fabriziosalmi/certmate#1207 ·
メンテナーはふだん 1 日以内に返信
-
[work-item] adam-adae-onset-emergence: document the minute-precision datetime coercion (ASTTMF not derivable)対応中かも @muse-yamaa-bot が今日担当しました。 オープンwork-item
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
CommunityToolkit/Aspire#2231 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
volcengine/OpenViking#5711 ·
メンテナーはふだん 1 日以内に返信