Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Per-request `_force_flush_otel()` in ADK template blocks streaming responses

Aperta
#6,807 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
68/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Tranquilla
Stack tecnologico
python

Direzione di ricerca

Inizia in vertexai/agent_engines/templates/adk.py, esamina streaming_agent_run_with_events, async_stream_query e _force_flush_otel, quindi riproduci una query in streaming con la telemetria abilitata e disabilitata. Il lavoro è completo quando la consegna della risposta non attende più un flush della telemetria per richiesta, mentre la telemetria continua a essere esportata in modo asincrono e la gestione della perdita di dati durante l’arresto rimane risolta.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

api: vertex-ai
Environment
  • google-cloud-aiplatform (vertexai SDK)
  • vertexai/agent_engines/templates/adk.py
  • Agent Runtime with identity_type=AGENT_IDENTITY
  • ADK >= 1.17.0
Description
Problem

The ADK template's streaming_agent_run_with_events and async_stream_query methods call _force_flush_otel() in their finally blocks on every request. This triggers a synchronous BatchSpanProcessor.force_flush() that blocks the response stream until all queued spans are exported to telemetry.googleapis.com.

In Agent Identity environments where the export uses certificate-bound tokens (mTLS via SPIFFE/WIF), the authentication overhead is significantly higher than standard OAuth, and the blocking duration becomes noticeable to end users. Setting GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY=false eliminates the overhead, confirming the flush is the sole cause.

Root Cause

_force_flush_otel() in adk.py calls tracer_provider.force_flush() via asyncio.to_thread(), which invokes BatchSpanProcessor._export(BatchExportStrategy.EXPORT_ALL) — a blocking synchronous export. The default OTEL_BSP_EXPORT_TIMEOUT is 30,000ms, so the flush can block for up to 30 seconds depending on the mTLS authentication latency.

The comment on the call site reads:

# Avoid telemetry data loss having to do with CPU throttling on instance turndown

This concern is valid, but per-request flush is the wrong mechanism for it. BatchSpanProcessor already exports asynchronously at a configurable interval (default 5s). Telemetry data loss on shutdown should be handled by a shutdown hook, not by blocking every request.

Reproduction
  1. Deploy an ADK agent to Agent Runtime with identity_type=AGENT_IDENTITY
  2. Ensure telemetry is enabled (GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY=true, or leave at default for ADK >= 1.17)
  3. Send a streaming query
  4. Observe a significant gap between agent execution completion and response delivery
  5. Set GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY=false and repeat — the gap disappears
Expected Behavior

The OTel flush should not block the request path. BatchSpanProcessor is designed to export spans asynchronously in background threads. Per-request force_flush() negates this design and introduces user-facing latency proportional to the telemetry export duration.

When running the same agent on Cloud Run with ADK Runner.run_async() directly (bypassing the ADK template wrapper), no such overhead exists — the agent execution time is the same, but the framework-level gap is negligible.

Code References

Call sites (per-request flush in finally blocks):

  • streaming_agent_run_with_events
  • async_stream_query

Flush implementation:

  • _force_flush_otel
Workaround

Set GOOGLE_CLOUD_AGENT_ENGINE_ENABLE_TELEMETRY=false in the Agent Runtime deployment's env_vars. This disables Cloud Trace export but does not affect BigQuery Agent Analytics Plugin (which uses its own data path).

Lingua principale
Python
Stelle
908
Fork
468
Merge medio
2g 4h
PR unite (30g)
25

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di googleapis/python-aiplatform

Tutte le issue di googleapis/python-aiplatform

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.