[Bug]: QueueShutDown spans still end as ERROR (and exceptions are recorded twice) because start_as_current_span sets status on exception
Les mainteneurs répondent en général sous 2 jours
Évaluation
Cette issue n'a pas encore été évaluée.
Description
What happened?
EventQueueSource.dequeue_event spans still end with status ERROR (AsyncQueueShutDown: ) on every normal queue teardown, even though #1075 (fix for #1065) added QueueShutDown to _NON_ERROR_EXCEPTIONS.
trace_function opens its span with OTel's defaults:
with tracer.start_as_current_span(actual_span_name, kind=kind) as span:
start_as_current_span defaults to record_exception=True and set_status_on_exception=True. So when the wrapper re-raises, OTel's use_span records the exception and calls span.set_status(ERROR, ...) for any Exception that leaves the block (opentelemetry/trace/__init__.py, use_span).
QueueShutDownstill ends as ERROR. Theexcept _NON_ERROR_EXCEPTIONSarm correctly leaves the status alone and re-raises. OTel then marks the span ERROR anyway.- Every
Exceptionis recorded twice. Bothexceptarms callspan.record_exception, and OTel records it a second time. So each such span carries twoexceptionevents.
CancelledError isn't affected, because it derives from BaseException and OTel only catches Exception.
The existing test (test_trace_function_async_non_error_exception_does_not_mark_span_error) uses a mocked tracer, so it only sees the wrapper's own set_status calls and not what OTel does on exit.
Expected: a span that ends with QueueShutDown is not ERROR, and a failing span has a single exception event.
Reproduction (a2a-sdk 1.2.1, opentelemetry-sdk 1.42.1, Python 3.12; same code on main at 6ff0a82):
import asyncio
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor
from opentelemetry.sdk.trace.export.in_memory_span_exporter import InMemorySpanExporter
exporter = InMemorySpanExporter()
provider = TracerProvider()
provider.add_span_processor(SimpleSpanProcessor(exporter))
trace.set_tracer_provider(provider)
from a2a.utils._async_queue_compat import QueueShutDown
from a2a.utils.telemetry import trace_function
@trace_function(span_name="shutdown")
async def shutdown():
raise QueueShutDown()
@trace_function(span_name="boom")
async def boom():
raise ValueError("real failure")
async def main():
for f in (shutdown, boom):
try:
await f()
except Exception:
pass
asyncio.run(main())
for s in exporter.get_finished_spans():
print(s.name, s.status.status_code.name, repr(s.status.description), [e.name for e in s.events])
Output:
shutdown ERROR 'AsyncQueueShutDown: ' ['exception', 'exception']
boom ERROR 'ValueError: real failure' ['exception', 'exception']
Both spans carry the exception twice, and boom's description is OTel's (ValueError: real failure), which overwrote the wrapper's str(e).
With the proposed fix, the same script prints:
shutdown UNSET None ['exception']
boom ERROR 'real failure' ['exception']
We noticed it on an A2A server running DefaultRequestHandlerV2: every streaming request leaves one ERROR dequeue_event span, which shows up as noise in the error rate of the trace backend.
Proposed fix: pass record_exception=False, set_status_on_exception=False to start_as_current_span in both wrappers, since the wrapper already does both itself. PR to follow.
Relevant log output
"name": "a2a.server.events.event_queue_v2.EventQueueSource.dequeue_event",
"kind": "SpanKind.SERVER",
"status": {
"status_code": "ERROR",
"description": "QueueShutDown: "
},
"events": [
{
"name": "exception",
"attributes": {
"exception.type": "culsans.QueueShutDown",
Code of Conduct
- I agree to follow this project's Code of Conduct
- Langage dominant
- Python
- Étoiles
- 2.2k
- Forks
- 509
- Merge moyen
- 3 j 18 h
- PR mergées (30 j)
- 44
Préparer son environnement
- Aucun Dockerfile ni fichier Docker Compose
- Propose un modèle de pull request
- Lire le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de a2aproject/a2a-python
-
[Bug]: REST task/request id sanitizationPeut-être pris @Linux2010 l’a pris il y a 103 jours. Ouvertemaintainers-only
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
a2aproject/a2a-python#805 · 1 commentaire ·
Les mainteneurs répondent en général sous 2 jours
-
[Feat]: Cluster mode: detect and recover tasks abandoned by a crashed replica (heartbeat/lease)Peut-être pris @rohityan l’a pris il y a 2 jours. Ouverte
a2aproject/a2a-python#1324 · 1 personne assignée ·
Les mainteneurs répondent en général sous 2 jours
-
[Bug]: Cluster mode: SubscribeToTask on a replica holding a paused ActiveTask returns a stale INPUT_REQUIRED snapshot and closesPeut-être pris @rohityan l’a pris il y a 2 jours. Ouvertecomponent: server
a2aproject/a2a-python#1323 · 1 personne assignée ·
Les mainteneurs répondent en général sous 2 jours
-
[Bug]: After a streamed task is cancelled, its background producer never finishesPeut-être pris @rohityan l’a pris il y a 2 jours. Ouvertecomponent: server status:awaiting response
a2aproject/a2a-python#1322 · 1 commentaire · 1 personne assignée ·
Les mainteneurs répondent en général sous 2 jours
-
v0.3 gRPC and REST SendMessage without configuration run non-blockingPeut-être pris @rohityan l’a pris il y a 2 jours. Ouvertequestion status:awaiting response
a2aproject/a2a-python#1321 · 1 commentaire · 1 personne assignée ·
Les mainteneurs répondent en général sous 2 jours
Toutes les issues de a2aproject/a2a-python
Issues similaires
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0Ouvertedocumentation priority: low size: S
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
Les mainteneurs répondent en général sous 1 jour
-
--csv-bom was never wired up: PR #850 added an unused helper parameter, so #846 is not fixedOuvertebug help wanted
Difficulté 2/5 1-3 heures Accessibilité débutants 75/100
Les mainteneurs répondent en général sous 1 jour
-
Broken link in index.rstOuvertedocumentation
Difficulté 1/5 Moins d'une heure Accessibilité débutants 65/100
ansys/pydpf-core#3547 ·
Les mainteneurs répondent en général sous 1 jour
-
core
Difficulté 2/5 1-3 heures Accessibilité débutants 70/100
vectorize-io/hindsight#5457 ·
Les mainteneurs répondent en général sous 1 jour
-
[Bug]: LangChain drops OpenAI Responses text blocks from session recordingPeut-être pris @ktz03 l’a pris aujourd’hui. Ouverte
Difficulté 2/5 1-3 heures Accessibilité débutants 72/100
volcengine/OpenViking#5806 ·
Les mainteneurs répondent en général sous 1 jour