auto_instrument() creates duplicate LLM spans for Google ADK Gemini calls
Los mantenedores suelen responder en 1 día
Evaluación
Este issue todavía no se ha evaluado.
Descripción
With braintrust.auto_instrument(), a Google ADK Gemini request produces two nested llm spans with identical token usage: ADK's llm_call [direct_response] and Google GenAI's generate_content.
Reproduction
Install:
pip install braintrust==0.45.0 google-adk==2.11.0 google-genai==2.28.0
Configure BRAINTRUST_API_KEY, BRAINTRUST_PROJECT, and GOOGLE_API_KEY through environment variables. Use a test project: this example makes a Gemini request and uploads its trace to Braintrust.
import asyncio
import os
import braintrust
from google.adk.agents import LlmAgent
from google.adk.runners import InMemoryRunner
from google.genai import types
logger = braintrust.init_logger(project=os.environ["BRAINTRUST_PROJECT"])
braintrust.auto_instrument()
async def main():
agent = LlmAgent(
name="hello",
model="gemini-2.5-flash",
instruction="Reply with one word.",
)
runner = InMemoryRunner(agent=agent, app_name="repro")
session = await runner.session_service.create_session(
app_name="repro", user_id="u"
)
msg = types.Content(
role="user", parts=[types.Part(text="Say hello.")]
)
with braintrust.start_span(name="adk_duplicate_span_repro"):
async for _ in runner.run_async(
user_id="u", session_id=session.id, new_message=msg
):
pass
asyncio.run(main())
logger.flush()
Observations and validation
Originally observed with Braintrust 0.41.0, Google ADK 1.14.1, and Google GenAI 1.75.0.
Also reproduced on October 7, 2026 using the latest stable versions available that day: Braintrust 0.45.0, Google ADK 2.11.0, and Google GenAI 2.28.0, on Python 3.13.12/macOS.
The latest-version test used the real ADK runner, Gemini adapter, GenAI SDK, and Braintrust instrumentation, with only the provider network response mocked. Spans were captured in memory without uploading. It did not verify live Gemini calls or Braintrust dashboard aggregation on these newer versions.
A synthetic response containing 23 prompt tokens and 62 completion tokens produced:
| Instrumentation | Provider requests | LLM spans | Summed prompt tokens | Summed completion tokens |
|---|---|---|---|---|
| Both enabled | 1 | 2 | 46 | 124 |
| ADK only | 1 | 1 | 23 | 62 |
| GenAI only | 1 | 1 | 23 | 62 |
Each configuration ran in a fresh process. Assertions confirmed that both spans carried identical usage and that generate_content was a direct child of llm_call [direct_response]. Dependency validation with pip check passed.
In the original live traces, trace summaries counted both spans and reported duplicated LLM-call and token totals. Exact token values can vary between live runs.
Expected behavior
Each underlying Gemini request should contribute usage and LLM-call count once, while retaining ADK agent/tool tracing and instrumentation of Gemini calls made outside ADK.
Suspected cause
_flow_call_llm_async_wrapper in py/src/braintrust/integrations/adk/tracing.py opens an llm span around BaseLlmFlow._call_llm_async and logs response usage.
The nested Google GenAI generate_content call is separately instrumented as an llm span and logs usage for the same request.
Impact and possible approach
Duplicated usage inflates token totals and can affect cost reporting based on those totals. Disabling ADK instrumentation loses its agent/tool tracing; disabling Google GenAI instrumentation loses automatic tracing of direct Gemini calls outside ADK in the same process.
One possible approach is to suppress the redundant provider span within an ADK model call, or retain it as a non-llm span without duplicated usage. Removing usage alone would leave the duplicate LLM-call count.
Related: #878, fixed by PR #880, concerns a different duplication mechanism involving wrap_openai() and auto_instrument().
- Lenguaje dominante
- Python
- Estrellas
- 21
- Forks
- 23
- Merge medio
- 18 h 23 min
- PR fusionados (30 d)
- 104
Preparar el entorno
- Incluye un Dockerfile o un archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de braintrustdata/braintrust-sdk-python
-
feature python
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
braintrustdata/braintrust-sdk-python#868 ·
Los mantenedores suelen responder en 1 día
-
new-integration python
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
braintrustdata/braintrust-sdk-python#854 ·
Los mantenedores suelen responder en 1 día
-
new-integration python
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
braintrustdata/braintrust-sdk-python#853 ·
Los mantenedores suelen responder en 1 día
-
new-integration python
Dificultad 4/5 3-5 días Aptitud para principiantes 52/100
braintrustdata/braintrust-sdk-python#852 ·
Los mantenedores suelen responder en 1 día
-
Migrate legacy HTTPConnection callers to the policy-aware Transport incrementallyPosiblemente ocupada @AbhiPrasad la tomó hace 9 días. Abiertopython
braintrustdata/braintrust-sdk-python#839 · 1 asignado ·
Los mantenedores suelen responder en 1 día
Todos los issues de braintrustdata/braintrust-sdk-python
Issues similares
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
topoteretes/cognee#5647 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
Sendspin/sendspin-python-cli#291 ·
Los mantenedores suelen responder en 6 días
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
awslabs/visual-asset-management-system#414 ·
Los mantenedores suelen responder en 1 día
-
bug v1 v2
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
modelcontextprotocol/python-sdk#3670 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
aicell-lab/bioengine#232 ·
Los mantenedores suelen responder en 1 día