[bot] OpenAI and Anthropic generation parameters not captured in span metadata
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 55/100
- Issue 类型
- 功能
- 描述清晰度
- 基本清楚
- 活跃度
- 冷清
- 技术栈
- java
调研方向
从 braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java 中的 tagOpenAIRequest 和 tagAnthropicRequest 开始,然后比较 Google GenAI 的 BraintrustApiClient.java handler 中的参数提取方式。检查现有的 OpenAI 和 Anthropic 测试文件,并为列出的生成参数和提供商特定配置添加覆盖。完成的标准是:一致地捕获受支持的请求元数据,并且测试能够对此进行断言。
由索引模型根据 Issue 内容生成。
描述
Summary
The OpenAI and Anthropic request taggers in InstrumentationSemConv only extract model and request routing info (request_path, request_base_uri, request_method) into span metadata. Generation parameters like temperature, max_tokens, tools, response_format, and provider-specific config like Anthropic's thinking are silently dropped.
In contrast, the Google GenAI handler in the same repo (BraintrustApiClient.tagSpan()) extracts temperature, topP, topK, maxOutputTokens, tools, toolConfig, safetySettings, responseMimeType, responseSchema, and more into metadata. This is an inconsistency within the repo — the Google GenAI handler provides materially more instrumentation detail for the same class of information.
What is missing
OpenAI — tagOpenAIRequest() (lines 78–108)
Currently captures only model in metadata. The following request parameters are silently dropped:
| Field | Purpose |
|---|---|
temperature |
Sampling temperature |
max_tokens / max_completion_tokens |
Output length limit |
top_p |
Nucleus sampling |
frequency_penalty, presence_penalty |
Repetition control |
tools |
Tool/function definitions |
response_format |
Structured output config (JSON mode, JSON Schema) |
reasoning_effort |
Reasoning effort for o-series models |
logprobs, top_logprobs |
Log probability settings |
stop |
Stop sequences |
For the Responses API, additional fields are missing: instructions, tools (with web_search, file_search, code_interpreter configs), reasoning (with effort and summary).
Anthropic — tagAnthropicRequest() (lines 163–205)
Currently captures only model in metadata. The following request parameters are silently dropped:
| Field | Purpose |
|---|---|
max_tokens |
Output length limit (required parameter) |
temperature |
Sampling temperature |
top_p, top_k |
Sampling parameters |
tools |
Tool definitions |
thinking |
Extended thinking config (type, budget_tokens) |
stop_sequences |
Stop sequences |
metadata |
User-provided metadata (e.g., user_id) |
The thinking config is particularly important: when a user enables extended thinking with a specific budget, this information is lost from the span. The Python SDK equivalent (issue braintrustdata/braintrust-sdk-python#107, now closed) captures thinking metadata.
Google GenAI — already captures these (for comparison)
BraintrustApiClient.tagSpan() (lines 64–95) extracts into metadata:
systemInstruction,tools,toolConfig,safetySettings,cachedContent- From
generationConfig:temperature,topP,topK,candidateCount,maxOutputTokens,stopSequences,responseMimeType,responseSchema
Impact
Users viewing traces in Braintrust can see generation parameters for Google GenAI calls but not for OpenAI or Anthropic calls. This makes it impossible to understand model configuration from the trace alone for the two most commonly used providers.
Braintrust docs status
- Braintrust tracing docs at https://www.braintrust.dev/docs/guides/tracing state that LLM spans show "the model, messages, parameters, token usage, and cost" — supported (parameters are documented as captured, but not implemented for OpenAI/Anthropic in this Java SDK)
- The Braintrust OpenAI docs mention temperature handling for GPT-5 models, implying it is a tracked parameter: unclear
Upstream sources
- OpenAI Chat Completions API: https://platform.openai.com/docs/api-reference/chat/create — documents all request parameters including
temperature,max_tokens,tools,response_format,reasoning_effort - OpenAI Responses API: https://platform.openai.com/docs/api-reference/responses/create — documents
instructions,tools,reasoning - Anthropic Messages API: https://docs.anthropic.com/en/api/messages — documents
max_tokens,temperature,tools,thinking,top_p,top_k,stop_sequences - Anthropic extended thinking: https://docs.anthropic.com/en/docs/build-with-claude/extended-thinking — documents
thinkingparameter withtypeandbudget_tokens
Local files inspected
braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java— lines 78–108 (tagOpenAIRequest: onlymodelextracted), lines 163–205 (tagAnthropicRequest: onlymodelextracted)braintrust-sdk/instrumentation/genai_1_18_0/src/main/java/com/google/genai/BraintrustApiClient.java— lines 64–95 (tagSpan: comprehensive parameter extraction into metadata)braintrust-sdk/instrumentation/openai_2_8_0/src/test/java/dev/braintrust/instrumentation/openai/v2_8_0/BraintrustOpenAITest.java— no test asserts generation parameters in metadatabraintrust-sdk/instrumentation/anthropic_2_2_0/src/test/java/dev/braintrust/instrumentation/anthropic/v2_2_0/BraintrustAnthropicTest.java— no test asserts generation parameters in metadata
- 主要语言
- Java
- 星标
- 21
- 派生
- 5
- 平均合并
- 2 天 34 分钟
- 30 天内合并 PR
- 6
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
braintrustdata/braintrust-sdk-java 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 82/100
-
难度 2/5 1-3 小时 新手友好度 82/100
-
难度 2/5 1-3 小时 新手友好度 82/100
-
难度 2/5 1-3 小时 新手友好度 78/100
-
难度 2/5 1-3 小时 新手友好度 68/100
查看 braintrustdata/braintrust-sdk-java 的全部 Issue
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 72/100
-
area/docs
难度 1/5 1 小时以内 新手友好度 88/100
维护者通常 1 天内回复
-
BoxChart rejects valid List.of data with NullPointerException可能已有人在做 @PHJ2000 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 76/100