Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[bot] OpenAI Batch API is not instrumented (mis-tagged as a generic LLM span)

已关闭
#164 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
48/100
Issue 类型
功能
描述清晰度
基本清楚
活跃度
冷清
技术栈
java

调研方向

从 braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java 开始,重点查看 tagOpenAIRequest()、tagOpenAIResponse() 和 getSpanName(),然后检查 OpenAI 模块中的 BraintrustOpenAI.java 和 TracingHttpClient.java。在现有的 instrumentation 测试中搜索可比的 provider batch 处理;当 Batch API 操作和可用的 batch 结果数据获得适当的 spans,并且 create、retrieve、list 和 cancel 都有 coverage 时,这项工作就完成了。

由索引模型根据 Issue 内容生成。

描述

Summary

The OpenAI instrumentation module (openai_2_15_0) generically intercepts every HTTP call via TracingHttpClient (swapped into ClientOptions.httpClient/originalHttpClient), so a call to client.batches().create(...) (or .retrieve()/.list()/.cancel()) does produce a span — but the shared tagging logic in InstrumentationSemConv has no awareness of the Batch API's request/response shape, so the span is actively mis-tagged rather than simply absent: it's marked span_attributes.type = "llm" as if it were a real model call, given a low-information span name ("batches"), and gets no model metadata, no input_json, and no output_json/metrics.

The OpenAI Batch API lets you submit up to 50,000 chat-completion/embeddings/responses/moderation requests as a single async job (POST /v1/batches); the job's eventual output file contains one JSONL line per request with the same generative output (token usage, model output) that this SDK already spans for synchronous calls — but none of that ever gets a Braintrust span today.

What is missing

In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:

  • getSpanName() (lines 518–529) switches on providerName + ":" + lastPathSegment. For POST /v1/batches, the last path segment is "batches", which matches neither the openai:completions nor openai:embeddings case, so it falls through to default -> lastSegment, yielding the literal span name "batches" instead of something descriptive like "openai.batches.create".
  • tagOpenAIRequest() (lines 110–141) unconditionally sets span_attributes = {"type":"llm"} (line 119) even though a batch-create call isn't itself a model invocation. It only reads metadata.model when requestJson.has("model") (line 129) and input_json from messages or an array-typed input (lines 133–137) — but a BatchCreateParams request body has none of these; it has input_file_id, endpoint (e.g. /v1/chat/completions), and completion_window. All of that is silently dropped.
  • tagOpenAIResponse() (lines 143–208) looks for choices or output for output_json (lines 149–153) and a top-level usage object for metrics (line 160) — a Batch object (returned by create/retrieve/list) has neither; it has id, status, output_file_id, error_file_id, request_counts, and timestamps. None of this is captured, and there is no instrumentation at all of retrieving/parsing the completed batch's output file (where the actual per-request custom_id + generative response.body results, including usage, become available) — so even a fully successful, completed batch job produces zero spans reflecting its actual generative work.
  • No test or example anywhere in the repo exercises client.batches() in any form (confirmed via grep for batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches).

Braintrust docs status: not_found

Checked https://www.braintrust.dev/docs/integrations/ai-providers/openai across all per-language sections (TypeScript, Python, Ruby, Go, Java, .NET): no mention of "batch" or "batches" anywhere. Its "What Braintrust traces" tables list only Chat Completion, Embedding, Moderation, openai.responses.create/parse/compact, Transcription, Translation, Speech, and Image Generation/Edit/Variation — the Batch API is absent for every language, not just Java. A broader site search only surfaces unrelated uses of "batch" (eval batches, UI batch labeling, batch-ingested span timestamps in the changelog).

Note: this repo's own gap-audit history already treats "Batch API not instrumented, mis-tagged as a generic LLM span" as a valid, in-scope finding — see the already-filed and still-open #155 for Anthropic's Message Batches API, which this issue mirrors for the OpenAI provider (a distinct upstream API/SDK, not a duplicate).

Upstream sources

Local repo files inspected

  • braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java — lines 110–141 (tagOpenAIRequest), 143–208 (tagOpenAIResponse), 518–529 (getSpanName)
  • braintrust-sdk/instrumentation/openai_2_15_0/src/main/java/dev/braintrust/instrumentation/openai/v2_15_0/BraintrustOpenAI.java and TracingHttpClient.java — generic transport-swap; produces a span for any OpenAI HTTP call including /v1/batches, with no batch-specific logic
  • Repo-wide grep for batch/Batch under braintrust-sdk/instrumentation/openai_2_15_0/ — zero matches (no test or example exercises this API)
主要语言
Java
星标
21
派生
5
平均合并
1 天 20 小时
30 天内合并 PR
8

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

braintrustdata/braintrust-sdk-java 的其他 Issue

查看 braintrustdata/braintrust-sdk-java 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。