Per-request `_force_flush_otel()` in ADK template blocks streaming responses
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 68/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Tranquilla
- Stack tecnologico
- python
- Ambito
- backend, observability-sre
Direzione di ricerca
Inizia in vertexai/agent_engines/templates/adk.py, esamina streaming_agent_run_with_events, async_stream_query e _force_flush_otel, quindi riproduci una query in streaming con la telemetria abilitata e disabilitata. Il lavoro è completo quando la consegna della risposta non attende più un flush della telemetria per richiesta, mentre la telemetria continua a essere esportata in modo asincrono e la gestione della perdita di dati durante l’arresto rimane risolta.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Environment
google-cloud-aiplatform(vertexai SDK)vertexai/agent_engines/templates/adk.py- Agent Runtime with
identity_type=AGENT_IDENTITY - ADK >= 1.17.0
Description
Problem
The ADK template's streaming_agent_run_with_events and async_stream_query methods call _force_flush_otel() in their finally blocks on every request. This triggers a synchronous BatchSpanProcessor.force_flush() that blocks the response stream until all queued spans are exported to telemetry.googleapis.com.
In Agent Identity environments where the export uses certificate-bound tokens (mTLS via SPIFFE/WIF), the authentication overhead is significantly higher than standard OAuth, and the blocking duration becomes noticeable to end users. Setting GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY=false eliminates the overhead, confirming the flush is the sole cause.
Root Cause
_force_flush_otel() in adk.py calls tracer_provider.force_flush() via asyncio.to_thread(), which invokes BatchSpanProcessor._export(BatchExportStrategy.EXPORT_ALL) — a blocking synchronous export. The default OTEL_BSP_EXPORT_TIMEOUT is 30,000ms, so the flush can block for up to 30 seconds depending on the mTLS authentication latency.
The comment on the call site reads:
# Avoid telemetry data loss having to do with CPU throttling on instance turndown
This concern is valid, but per-request flush is the wrong mechanism for it. BatchSpanProcessor already exports asynchronously at a configurable interval (default 5s). Telemetry data loss on shutdown should be handled by a shutdown hook, not by blocking every request.
Reproduction
- Deploy an ADK agent to Agent Runtime with
identity_type=AGENT_IDENTITY - Ensure telemetry is enabled (
GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY=true, or leave at default for ADK >= 1.17) - Send a streaming query
- Observe a significant gap between agent execution completion and response delivery
- Set
GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY=falseand repeat — the gap disappears
Expected Behavior
The OTel flush should not block the request path. BatchSpanProcessor is designed to export spans asynchronously in background threads. Per-request force_flush() negates this design and introduces user-facing latency proportional to the telemetry export duration.
When running the same agent on Cloud Run with ADK Runner.run_async() directly (bypassing the ADK template wrapper), no such overhead exists — the agent execution time is the same, but the framework-level gap is negligible.
Code References
Call sites (per-request flush in finally blocks):
streaming_agent_run_with_eventsasync_stream_query
Flush implementation:
_force_flush_otel
Workaround
Set GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY=false in the Agent Runtime deployment's env_vars. This disables Cloud Trace export but does not affect BigQuery Agent Analytics Plugin (which uses its own data path).
- Lingua principale
- Python
- Stelle
- 908
- Fork
- 468
- Merge medio
- 2g 4h
- PR unite (30g)
- 25
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di googleapis/python-aiplatform
-
Protobuf 7.35.1+ supportApertaapi: vertex-ai
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 85/100
googleapis/python-aiplatform#7132 · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
api: vertex-ai
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
googleapis/python-aiplatform#7097 ·
I maintainer di solito rispondono entro 1 giorno
-
CustomContainerTrainingJob.run drops max_wait_duration=0 instead of requesting indefinite DWS waitApertaapi: vertex-ai
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
googleapis/python-aiplatform#7067 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
api: vertex-ai
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
googleapis/python-aiplatform#6877 ·
I maintainer di solito rispondono entro 1 giorno
-
api: vertex-ai
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
googleapis/python-aiplatform#6865 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di googleapis/python-aiplatform
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
mikf/gallery-dl#9791 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
fossasia/eventyay#6151 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
P4: low tooling
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
jeffknupp/association#318 ·
-
azure-cost bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
microsoft/GitHub-Copilot-for-Azure#3330 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
raullenchai/Rapid-MLX#4097 ·
I maintainer di solito rispondono entro 1 giorno