Hacktoberfest 2026 : les issues que les mainteneurs ont marquées pour octobre, ouvertes et accessibles aux débutants. Parcourir les issues Hacktoberfest

[Bug]: QueueShutDown spans still end as ERROR (and exceptions are recorded twice) because start_as_current_span sets status on exception

Fermée
#1,308 0 commentaires 0 réactions 2 personnes assignées Voir sur GitHub

Les mainteneurs répondent en général sous 2 jours

@rohityan y travaille déjà.

Depuis le 3/10/2026.

  • #1309 par @ferponse — ouverte

Évaluation

Cette issue n'a pas encore été évaluée.

Description

What happened?

EventQueueSource.dequeue_event spans still end with status ERROR (AsyncQueueShutDown: ) on every normal queue teardown, even though #1075 (fix for #1065) added QueueShutDown to _NON_ERROR_EXCEPTIONS.

trace_function opens its span with OTel's defaults:

with tracer.start_as_current_span(actual_span_name, kind=kind) as span:

start_as_current_span defaults to record_exception=True and set_status_on_exception=True. So when the wrapper re-raises, OTel's use_span records the exception and calls span.set_status(ERROR, ...) for any Exception that leaves the block (opentelemetry/trace/__init__.py, use_span).

  • QueueShutDown still ends as ERROR. The except _NON_ERROR_EXCEPTIONS arm correctly leaves the status alone and re-raises. OTel then marks the span ERROR anyway.
  • Every Exception is recorded twice. Both except arms call span.record_exception, and OTel records it a second time. So each such span carries two exception events.

CancelledError isn't affected, because it derives from BaseException and OTel only catches Exception.

The existing test (test_trace_function_async_non_error_exception_does_not_mark_span_error) uses a mocked tracer, so it only sees the wrapper's own set_status calls and not what OTel does on exit.

Expected: a span that ends with QueueShutDown is not ERROR, and a failing span has a single exception event.

Reproduction (a2a-sdk 1.2.1, opentelemetry-sdk 1.42.1, Python 3.12; same code on main at 6ff0a82):

import asyncio
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter

exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)

from a2a.utils._async_queue_compat import QueueShutDown
from a2a.utils.telemetry import trace_function

@trace_function(span_name="shutdown")
async def shutdown():
    raise QueueShutDown()

@trace_function(span_name="boom")
async def boom():
    raise ValueError("real failure")

async def main():
    for f in (shutdown, boom):
        try:
            await f()
        except Exception:
            pass

asyncio.run(main())
for s in exporter.get_finished_spans():
    print(s.name, s.status.status_code.name, repr(s.status.description), [e.name for e in s.events])

Output:

shutdown ERROR 'AsyncQueueShutDown: ' ['exception', 'exception']
boom ERROR 'ValueError: real failure' ['exception', 'exception']

Both spans carry the exception twice, and boom's description is OTel's (ValueError: real failure), which overwrote the wrapper's str(e).

With the proposed fix, the same script prints:

shutdown UNSET None ['exception']
boom ERROR 'real failure' ['exception']

We noticed it on an A2A server running DefaultRequestHandlerV2: every streaming request leaves one ERROR dequeue_event span, which shows up as noise in the error rate of the trace backend.

Proposed fix: pass record_exception=False, set_status_on_exception=False to start_as_current_span in both wrappers, since the wrapper already does both itself. PR to follow.

Relevant log output
"name": "a2a.server.events.event_queue_v2.EventQueueSource.dequeue_event",
"kind": "SpanKind.SERVER",
"status": {
    "status_code": "ERROR",
    "description": "QueueShutDown: "
},
"events": [
    {
        "name": "exception",
        "attributes": {
            "exception.type": "culsans.QueueShutDown",
Code of Conduct
  • I agree to follow this project's Code of Conduct
Langage dominant
Python
Étoiles
2.2k
Forks
509
Merge moyen
3 j 18 h
PR mergées (30 j)
44

Préparer son environnement

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de a2aproject/a2a-python

Toutes les issues de a2aproject/a2a-python

Issues similaires

Plus d'issues Python

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.