Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Responses.retrieve is not instrumented, so background responses lose output, usage and cost

Abierto
#1,834 1 comentario 0 reacciones 1 asignado Ver en GitHub

@hassiebp ya está trabajando en esto.

Desde el 24/8/2026.

Evaluación

Este issue todavía no se ha evaluado.

Descripción

billing bug feat-billing feat-llm-cost-tracking integration-openai sdk-python

Summary

Responses.create and Responses.parse are instrumented, but Responses.retrieve is not. For background responses that is the call that carries the actual result, so the generation in the trace never receives its output, token counts or cost.

Why this path matters

With background=True the API returns immediately with status="queued" and no usage. The result is collected later, either with client.responses.retrieve(response_id) or by resuming the stream with client.responses.stream(response_id=...).

The second one is easy to miss: stream() has two branches. Creating a new response goes through partial(self.create, ..., stream=True) and is traced correctly. Resuming an existing one goes through self.retrieve(response_id=..., stream=True) and is not.

Reproduction

No API key needed — this shows which method each path resolves to:

import openai, langfuse.openai
from openai.resources.responses import Responses

calls = []
Responses.create = lambda self, **kw: calls.append(("create", kw.get("stream"))) or iter(())
Responses.retrieve = lambda self, **kw: calls.append(("retrieve", kw.get("stream"))) or iter(())

c = openai.OpenAI(api_key="sk-test", base_url="http://localhost:9")

with c.responses.stream(model="gpt-5.6", input="hi"):   # new response
    pass
with c.responses.stream(response_id="resp_123"):        # resumed response
    pass

print(calls)  # [('create', True), ('retrieve', True)]

And which of those the SDK patches:

import openai
from openai.resources.responses import Responses
before = {m: getattr(Responses, m) for m in ("create", "parse", "retrieve")}
import langfuse.openai
for m in before:
    print(m, before[m] is not getattr(Responses, m))

# create   True
# parse    True
# retrieve False

grep -c background langfuse/openai.py returns 0, and tests/unit/test_openai.py covers neither retrieve nor .stream(.

Impact

A background response produces a generation with an input and a model, and then nothing else: no output, no usage_details, no cost. On a dashboard that reads as a call that was free, which is worse than a call that is missing — the trace looks complete.

Proposed fix

Add Responses.retrieve / AsyncResponses.retrieve to OPENAI_METHODS_V1. retrieve() carries no prompt of its own, so the response id can serve as the generation input; model, output and usage are already extracted from the response object, and the existing streaming extraction handles stream=True.

One open question for maintainers: retrieving a response that was already traced at creation time (non-background usage) would then record its usage a second time. Options are to instrument every retrieve, or only the stream=True path, or to skip usage when the response was not queued. Happy to follow whichever you prefer — PR attached implements the straightforward version.

Version info

  • langfuse: main (4.x)
  • openai: 2.29.0
  • Python: 3.14
Lenguaje dominante
Python
Estrellas
468
Forks
349
Merge medio
15 h 51 min
PR fusionados (30 d)
24

Guía de contribución

Abrir la guía de contribución

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de langfuse/langfuse-python

Todos los issues de langfuse/langfuse-python

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.