[Bug]: QueueShutDown spans still end as ERROR (and exceptions are recorded twice) because start_as_current_span sets status on exception
Maintainer thường phản hồi trong vòng 2 ngày
Đánh giá
Issue này chưa được đánh giá.
Mô tả
What happened?
EventQueueSource.dequeue_event spans still end with status ERROR (AsyncQueueShutDown: ) on every normal queue teardown, even though #1075 (fix for #1065) added QueueShutDown to _NON_ERROR_EXCEPTIONS.
trace_function opens its span with OTel's defaults:
with tracer.start_as_current_span(actual_span_name, kind=kind) as span:
start_as_current_span defaults to record_exception=True and set_status_on_exception=True. So when the wrapper re-raises, OTel's use_span records the exception and calls span.set_status(ERROR, ...) for any Exception that leaves the block (opentelemetry/trace/__init__.py, use_span).
QueueShutDownstill ends as ERROR. Theexcept _NON_ERROR_EXCEPTIONSarm correctly leaves the status alone and re-raises. OTel then marks the span ERROR anyway.- Every
Exceptionis recorded twice. Bothexceptarms callspan.record_exception, and OTel records it a second time. So each such span carries twoexceptionevents.
CancelledError isn't affected, because it derives from BaseException and OTel only catches Exception.
The existing test (test_trace_function_async_non_error_exception_does_not_mark_span_error) uses a mocked tracer, so it only sees the wrapper's own set_status calls and not what OTel does on exit.
Expected: a span that ends with QueueShutDown is not ERROR, and a failing span has a single exception event.
Reproduction (a2a-sdk 1.2.1, opentelemetry-sdk 1.42.1, Python 3.12; same code on main at 6ff0a82):
import asyncio
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
from a2a.utils._async_queue_compat import QueueShutDown
from a2a.utils.telemetry import trace_function
@trace_function(span_name="shutdown")
async def shutdown():
raise QueueShutDown()
@trace_function(span_name="boom")
async def boom():
raise ValueError("real failure")
async def main():
for f in (shutdown, boom):
try:
await f()
except Exception:
pass
asyncio.run(main())
for s in exporter.get_finished_spans():
print(s.name, s.status.status_code.name, repr(s.status.description), [e.name for e in s.events])
Output:
shutdown ERROR 'AsyncQueueShutDown: ' ['exception', 'exception']
boom ERROR 'ValueError: real failure' ['exception', 'exception']
Both spans carry the exception twice, and boom's description is OTel's (ValueError: real failure), which overwrote the wrapper's str(e).
With the proposed fix, the same script prints:
shutdown UNSET None ['exception']
boom ERROR 'real failure' ['exception']
We noticed it on an A2A server running DefaultRequestHandlerV2: every streaming request leaves one ERROR dequeue_event span, which shows up as noise in the error rate of the trace backend.
Proposed fix: pass record_exception=False, set_status_on_exception=False to start_as_current_span in both wrappers, since the wrapper already does both itself. PR to follow.
Relevant log output
"name": "a2a.server.events.event_queue_v2.EventQueueSource.dequeue_event",
"kind": "SpanKind.SERVER",
"status": {
"status_code": "ERROR",
"description": "QueueShutDown: "
},
"events": [
{
"name": "exception",
"attributes": {
"exception.type": "culsans.QueueShutDown",
Code of Conduct
- I agree to follow this project's Code of Conduct
- Ngôn ngữ chính
- Python
- Star
- 2.2k
- Fork
- 499
- Merge trung bình
- 3 ngày 13 giờ
- Pull request đã merge (30 ngày)
- 40
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của a2aproject/a2a-python
-
[Bug]: REST task/request id sanitizationCó thể đã có người làm @Linux2010 đã nhận 100 ngày trước. Đang mởmaintainers-only
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
a2aproject/a2a-python#805 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug]: DefaultRequestHandler runs follow-up messages on a task with the first request's contextvarsCó thể đã có người làm @rohityan đã nhận hôm nay. Đang mở
a2aproject/a2a-python#1316 · 1 reaction · 2 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug]: Push notification store failure rewrites a completed task as FAILED (DefaultRequestHandlerV2)Có thể đã có người làm @rohityan đã nhận 1 ngày trước. Đang mở
a2aproject/a2a-python#1313 · 2 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
v0.3 JSON-RPC and REST GetTask return no history when history_length is 0Có thể đã có người làm @rohityan đã nhận 1 ngày trước. Đang mở
a2aproject/a2a-python#1311 · 2 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
-
[Bug]: DefaultRequestHandlerV2 keeps the ActiveTask (producer, consumer, 2 dispatchers) alive forever after a direct Message or input-required responseCó thể đã có người làm @rohityan đã nhận 3 ngày trước. Đang mở
a2aproject/a2a-python#1296 · 2 người được giao ·
Maintainer thường phản hồi trong vòng 2 ngày
Tất cả issue của a2aproject/a2a-python
Issue tương tự
-
needs-human needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
gke-labs/kube-agents#2400 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Device Details tables: FS/SF columns contradict each other (nfet_01v8 Vt row, pfet_01v8 Idsat row)Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
google/skywater-pdk#450 ·
-
Drained trajectory arrays are overwritten when the sequence buffer is reusedCó thể đã có người làm @sylvesterkaczmarek đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
google-deepmind/bsuite#56 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
LearningCircuit/local-deep-research#7206 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[TASK] Document technology stackĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
chingu-voyages/V62-tier3-team-33#285 ·
Maintainer thường phản hồi trong vòng 1 ngày