Responses.retrieve is not instrumented, so background responses lose output, usage and cost
维护者通常 1 天内回复
@hassiebp 已经在做这个了。
开始于 2026年8月24日。
评估
这个 Issue 还没有评估数据。
描述
Summary
Responses.create and Responses.parse are instrumented, but Responses.retrieve is not. For background responses that is the call that carries the actual result, so the generation in the trace never receives its output, token counts or cost.
Why this path matters
With background=True the API returns immediately with status="queued" and no usage. The result is collected later, either with client.responses.retrieve(response_id) or by resuming the stream with client.responses.stream(response_id=...).
The second one is easy to miss: stream() has two branches. Creating a new response goes through partial(self.create, ..., stream=True) and is traced correctly. Resuming an existing one goes through self.retrieve(response_id=..., stream=True) and is not.
Reproduction
No API key needed — this shows which method each path resolves to:
import openai, langfuse.openai
from openai.resources.responses import Responses
calls = []
Responses.create = lambda self, **kw: calls.append(("create", kw.get("stream"))) or iter(())
Responses.retrieve = lambda self, **kw: calls.append(("retrieve", kw.get("stream"))) or iter(())
c = openai.OpenAI(api_key="sk-test", base_url="http://localhost:9")
with c.responses.stream(model="gpt-5.6", input="hi"): # new response
pass
with c.responses.stream(response_id="resp_123"): # resumed response
pass
print(calls) # [('create', True), ('retrieve', True)]
And which of those the SDK patches:
import openai
from openai.resources.responses import Responses
before = {m: getattr(Responses, m) for m in ("create", "parse", "retrieve")}
import langfuse.openai
for m in before:
print(m, before[m] is not getattr(Responses, m))
# create True
# parse True
# retrieve False
grep -c background langfuse/openai.py returns 0, and tests/unit/test_openai.py covers neither retrieve nor .stream(.
Impact
A background response produces a generation with an input and a model, and then nothing else: no output, no usage_details, no cost. On a dashboard that reads as a call that was free, which is worse than a call that is missing — the trace looks complete.
Proposed fix
Add Responses.retrieve / AsyncResponses.retrieve to OPENAI_METHODS_V1. retrieve() carries no prompt of its own, so the response id can serve as the generation input; model, output and usage are already extracted from the response object, and the existing streaming extraction handles stream=True.
One open question for maintainers: retrieving a response that was already traced at creation time (non-background usage) would then record its usage a second time. Options are to instrument every retrieve, or only the stream=True path, or to skip usage when the response was not queued. Happy to follow whichever you prefer — PR attached implements the straightforward version.
Version info
- langfuse: main (4.x)
- openai: 2.29.0
- Python: 3.14
- 主要语言
- Python
- 星标
- 495
- 派生
- 358
- 平均合并
- 16 小时 54 分钟
- 30 天内合并 PR
- 20
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
langfuse/langfuse-python 的其他 Issue
-
bug: GCS upload detection in MediaManager uses substring match on the full URL可能已有人在做 @hassiebp 于 1 天前认领。 未关闭bug feat-multimodal-media sdk-python
langfuse/langfuse-python#1913 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
bug: get_dataset_run / get_dataset_runs / delete_dataset_run are unusable on Langfuse v4可能已有人在做 @hassiebp 于 5 天前认领。 未关闭bug feat-datasets feat-experiments sdk-python
langfuse/langfuse-python#1906 · 已指派 1 人 ·
维护者通常 1 天内回复
-
mask is not applied to create_dataset_item or create_score(comment=)可能已有人在做 @hassiebp 于 9 天前认领。 未关闭bug compliance feat-data-masking sdk-python security
langfuse/langfuse-python#1896 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
Scores bypass sample_rate since v4可能已有人在做 @hassiebp 于 14 天前认领。 未关闭feat-scores sdk-python unconfirmed-bug
langfuse/langfuse-python#1890 · 已指派 1 人 ·
维护者通常 1 天内回复
-
batch_evaluation fails on self-hosted v4 events_only deployments (uses unavailable v3 read endpoints)可能已有人在做 @hassiebp 于 27 天前认领。 未关闭bug feat-evals sdk-python
langfuse/langfuse-python#1861 · 2 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
查看 langfuse/langfuse-python 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 88/100
BasedHardware/omi#20271 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 92/100
openai/openai-cookbook#3153 ·
维护者通常 1 天内回复
-
cvss-severity:high devguard l3montree-cybersecurity/devguard/devguard pkg:golang/github.com/l3montree-dev/devguard risk:low state:open
难度 2/5 1-3 小时 新手友好度 65/100
l3montree-dev/devguard#3146 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 76/100
-
bug confirmed issue
难度 2/5 1-3 小时 新手友好度 76/100
open-webui/open-webui#31849 · 2 条评论 ·
维护者通常 1 天内回复