[Bug]: some observability capabilities loss on A2A
Maintainer antworten meist innerhalb von 1 Tag
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 35/100
- Issue-Typ
- Bug
- Klarheit
- Klar beschrieben
- Aktivitätsstatus
- Aktiv
- Bereich
- observability-sre
Rechercherichtung
Die Problembeschreibung nennt mehrere Lücken beim Tracing, bei Metriken und beim Logging im a2a-java SDK. Untersuche zunächst die Span-Erstellung im core module, insbesondere beim SSE-Streaming im Server-Agenten und bei der Zustellung von Push-Benachrichtigungen. Prüfe, an welchen Stellen gen_ai.agent.a2a.*-Attribute anstelle der offiziellen gen_ai.*-Semantikkonventionen gesetzt werden. Suche nach einer bereits vorhandenen Instrumentierung für Metriken. „Done“ bedeutet, die Architektur so weit zu verstehen, dass konkrete Codeänderungen für mindestens eines der acht aufgeführten Issues vorgeschlagen werden können.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
What happened?
First, I would like to confirm whether the community has a roadmap for observability capabilities and whether the current capabilities are considered complete by the community.
Below are some issues I encountered when building observability capabilities based on a2a-java.
Issue 1: Server-side SSE per-event spans are missing
Summary
When a Server Agent sends SSE streaming events back to the Client, there are no per-event spans on the Server side. The Server span ends as soon as the stream publisher is created ("Stream publisher created"), making the actual event emission process a complete black box.
Current Behavior
Client Agent Server Agent
───────── ─────────
[CLIENT Span: sendMessage] ──HTTP──→ [SERVER Span: onMessageSendStream]
│ │
│ Span ends after stream init │ Span ends after publisher returns
▼ ▼
[CLIENT Span: event-1] ❌ No Span
[CLIENT Span: event-2] ❌ No Span
[CLIENT Span: event-3] ❌ No Span
│ │
▼ ▼
Link to initial Event sending completely
CLIENT Span (non-parent-child) untraceable
- The Server span closes before any SSE events are actually sent.
- The Client side has per-event spans (linked via
addLink), but the Server side has nothing. - Client and Server event spans cannot be correlated one-to-one.
Expected Behavior
Each SSE event sent by the Server should have an INTERNAL span as a child of the Server span (via context restoration), with attributes such as streaming.event.kind (status / artifact / completed / error). This would close the end-to-end trace loop.
Impact
- Server-side event emission is a black box — no visibility into which events were sent, when, or how long each took.
- End-to-end trace tree is broken: the Server span has no children, and Client event spans are only linked (not parent-child), so they don't appear in the trace tree view.
- Cannot correlate Client-side event spans with Server-side event sending.
Issue 2: Push notification delivery is completely untraceable
Summary
Push notification delivery has no tracing at all. The active push callback from the Server Agent to the Client's webhook endpoint cannot be correlated to the original request.
Current Behavior
- Push notification configuration management (set/get/delete webhook) has spans.
- But the actual push notification delivery (the webhook callback) has zero tracing.
- There is no span, no trace context propagation, and no way to link the push delivery back to the original task/request.
Expected Behavior
- Push notification delivery should create a span with attributes such as
push.notification.url,notification.topic, andstreaming.event.kind. - The delivery span should be linkable to the original request span (via stored context or span link).
Impact
- Cannot track whether push notifications were successfully delivered.
- Cannot measure push notification delivery latency.
- Cannot correlate push notifications back to the original task or subscription.
Issue 3: No business-level semantic attributes on spans
Summary
A2A protocol-layer spans carry basic attributes (operation name, task/message/context ID), but there are no business-level semantic attributes such as task type,.
Current Behavior
no genai attributes, no business attributes.
Expected Behavior
The SDK should provide a mechanism (e.g., callback, annotation, or API) for users to enrich spans with these business-level attributes.
Impact
- Traces show "what happened at the protocol level" but not "what happened at the business level."
Issue 4: Task state transitions are not tracked
Summary
Task status attributes (e.g., submitted → working → completed) are reserved in the code but not activated. Task state transitions are invisible in traces.
Current Behavior
- The span attribute for task status appears to be defined but never populated with actual values.
- When a task transitions from
submittedtoworkingtocompleted, no span attribute records this transition.
Expected Behavior
Each SSE status event should carry a task.status attribute reflecting the current task state, enabling state-machine visualization and duration analysis per state.
Impact
- Cannot track task progress through its lifecycle.
- Cannot measure time spent in each task state.
- Cannot build task state-machine visualizations in observability dashboards.
Issue 5: No Metrics implementation in core module
Summary
The a2a-java core module has zero Metrics implementation. There are no histograms, counters, or any other instruments measuring operation duration, throughput, or token usage.
Current Behavior
| Expected Metric | Type | Status |
|---|---|---|
gen_ai.client.operation.duration |
Histogram | ❌ Missing |
gen_ai.client.token.usage |
Histogram | ❌ Missing |
a2a.task.request.duration |
Histogram | ❌ Missing |
- There is no
gen_ai.client.operation.durationmetric, which is the standard OTel GenAI operation duration histogram.
Expected Behavior
The core SDK module should provide built-in metrics instruments, at minimum:
gen_ai.client.operation.duration(OTel GenAI standard) — operation latency histogramgen_ai.client.token.usage(OTel GenAI standard) — token usage histogram bygen_ai.token.type- Task request duration covering the full SSE streaming lifecycle (from stream initiation to last event)
Impact
- Cannot measure operation latency or throughput at scale.
- Cannot set up SLO/SLA alerts based on latency percentiles.
- Cannot quantify agent interaction performance trends.
- Token usage is not tracked, making cost analysis impossible.
Issue 6: Span attributes use custom gen_ai.agent.a2a.* prefix instead of official gen_ai.* SemConv
Summary
All A2A span attributes use a custom gen_ai.agent.a2a.* prefix (e.g., gen_ai.agent.a2a.operation.name, gen_ai.agent.a2a.context_id) instead of the official OpenTelemetry GenAI Semantic Conventions (gen_ai.*). This prevents automatic integration with OTel GenAI dashboards and alerting rules.
Current Behavior
| Official GenAI SemConv Attribute | a2a-java Attribute | Status |
|---|---|---|
gen_ai.provider.name (formerly gen_ai.system, deprecated) |
— | ❌ Missing |
gen_ai.operation.name |
gen_ai.agent.a2a.operation.name |
⚠️ Prefixed |
gen_ai.conversation.id |
gen_ai.agent.a2a.context_id |
⚠️ Prefixed (code comment notes should use standard key) |
gen_ai.request.model |
— | ❌ Missing |
gen_ai.response.model |
— | ❌ Missing |
gen_ai.usage.input_tokens |
— | ❌ Missing |
gen_ai.usage.output_tokens |
— | ❌ Missing |
gen_ai.client.operation.duration |
— | ❌ Missing |
gen_ai.client.token.usage |
— | ❌ Missing |
server.address / server.port |
— | ❌ Missing |
Notably, the sample/example code uses the correct gen_ai.* keys for metric attributes, creating an inconsistency between the core module and the examples.
Expected Behavior
- Use official
gen_ai.*attribute keys as defined by the OTel GenAI SemConv. - For A2A-specific extensions that don't have an official SemConv equivalent, use a clearly namespaced prefix (e.g.,
a2a.*) rather than embedding undergen_ai.agent.a2a.*. - Align the core module attributes with the example code (which already uses correct keys).
Impact
- OTel GenAI-compatible dashboards (e.g., Grafana, Jaeger) cannot auto-discover and display A2A spans.
- Alerting rules built for
gen_ai.*attributes do not match A2A traces. - Users must build custom queries and dashboards instead of reusing existing GenAI observability tooling.
- Inconsistency between core module and example code creates confusion.
Issue 7: SSE streaming operation duration is incomplete — span ends at stream initiation
Summary
The current span for SSE streaming requests ends immediately after the stream is initiated (when the publisher returns), not when the stream actually completes. This means gen_ai.client.operation.duration (if it existed) would only record the stream setup time — potentially 0ms — rather than the full end-to-end streaming duration.
Current Behavior
CLIENT Span: sendMessage
│
│ ← Span ends here (stream publisher created)
│
[CLIENT Span: event-1] (addLink)
[CLIENT Span: event-2] (addLink)
[CLIENT Span: event-3] (addLink)
│
│ ← Stream actually ends here (no span records this)
▼
- The parent span closes before any events are sent.
- There is no span or metric that captures the full duration from stream initiation to the last event.
- The measured "operation duration" may be 0ms or near-0ms, completely misleading.
Expected Behavior
- Either extend the parent span's lifetime to cover the full streaming duration, or
- Record a dedicated metric (
a2a.task.request.durationor similar) that measures the time from stream initiation to the final event.
Impact
- Operation latency appears as ~0ms in dashboards, masking actual processing time.
- Cannot measure true end-to-end streaming duration.
- SLO/SLA monitoring based on operation duration would be incorrect.
Issue 8: No structured payload logging
Summary
There is no structured logging of A2A protocol message payloads. The only option is an opt-in toString() on span attributes, which is not a proper logging mechanism.
Current Behavior
- a2a-java has
gen_ai.agent.a2a.request/gen_ai.agent.a2a.responsespan attributes (enabled via system property), but these are span attributes, not structured logs. - There is no structured log event for task requests, state changes, result artifacts, negotiation payloads, authorization decisions, or notification deliveries.
- No payload size truncation or sensitive data masking is provided.
Expected Behavior
- Provide structured log events (at DEBUG level by default) for key protocol messages: task requests, state changes, result artifacts, negotiation payloads, authorization decisions, notification subscriptions and deliveries.
- Include configurable payload size truncation (default 4KB) with
[truncated]marker. - Provide sensitive data masking with customizable strategies.
- Logs should be correlated to traces via
trace_id/span_id.
Impact
- Cannot inspect actual message payloads when debugging issues.
- No audit trail of authorization decisions or negotiation content.
- Cannot correlate log entries to traces for root-cause analysis.
- The
toString()approach is not suitable for production use (no truncation, no masking, no structured fields).
Summary Table
| # | Issue | Severity | Category |
|---|---|---|---|
| 1 | Server-side SSE per-event spans missing | High | Tracing Gap |
| 2 | Push notification delivery untraceable | High | Tracing Gap |
| 3 | No business-level semantic attributes | High | Semantic Gap |
| 4 | Task state transitions not tracked | Medium | Tracing Gap |
| 5 | No Metrics implementation in core module | High | Metrics Gap |
| 6 | Custom gen_ai.agent.a2a.* prefix vs official gen_ai.* SemConv |
High | SemConv Misalignment |
| 7 | SSE streaming duration incomplete (span ends at initiation) | High | Tracing Gap |
| 8 | No structured payload logging | Medium | Logging Gap |
Relevant log output
Code of Conduct
- I agree to follow this project's Code of Conduct
- Vorherrschende Sprache
- Java
- Sterne
- 500
- Forks
- 179
- Ø Merge
- 1 T. 16 Std.
- Gemergte PRs (30 T.)
- 48
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus a2aproject/a2a-java
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 75/100
a2aproject/a2a-java#1012 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 65/100
a2aproject/a2a-java#464 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
Create a sample that demonstrates how to use a shared contextId across multiple tasksEvtl. wieder frei @tanish111 hat das vor 344 Tagen übernommen, und es ist kein Pull Request offen. Offensample
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 70/100
a2aproject/a2a-java#375 · 5 Kommentare · 1 zugewiesene Person ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 55/100
a2aproject/a2a-java#1161 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 3/5 1-2 Tage Anfängerfreundlichkeit 65/100
a2aproject/a2a-java#1121 · 2 Kommentare ·
Maintainer antworten meist innerhalb von 1 Tag
Alle Issues in a2aproject/a2a-java
Ähnliche Issues
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 84/100
beehive-lab/TornadoVM#1151 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 77/100
FasterXML/jackson-dataformats-binary#823 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 90/100
Maintainer antworten meist innerhalb von 1 Tag
-
bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
Maintainer antworten meist innerhalb von 1 Tag