[bot] Anthropic Message Batches API is not instrumented (mis-tagged as a generic LLM span)
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 48/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- java
- Lĩnh vực
- api, observability
Hướng nghiên cứu
Bắt đầu với InstrumentationSemConv.java, đặc biệt là tagAnthropicRequest(), tagAnthropicResponse() và getSpanName(), sau đó kiểm tra TracingHttpClient.java và ContextCapturingProxy.java để hiểu luồng span hiện có. Xem lại BraintrustAnthropicTest.java và các cấu trúc request cũng như kết quả batch của Anthropic trong tài liệu upstream được liên kết. Công việc được xem là hoàn tất khi các thao tác batch có span và metadata có ý nghĩa, đồng thời coverage của resultsStreaming() xác minh các kết quả thành công và lỗi cho từng request.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
The Anthropic instrumentation module (anthropic_2_2_0) only has shape-aware tagging for the Messages API (messages().create/streaming). It has no support at all for the Message Batches API (client.messages().batches()), a stable, GA, actively-developed part of the Anthropic Java SDK for submitting many message-generation requests as a single asynchronous job.
Because the generic transport-swap (TracingHttpClient) still intercepts the batch HTTP calls, a batches().create(...) call does produce a span — but it is actively mis-tagged rather than simply absent: it gets span_attributes.type = "llm" (as if it were a real model call), a low-information span name ("batches", from the path-segment fallback), no model in metadata, and no input_json at all — because the request/response shapes for batches are structurally different from a single Messages call and none of the existing field checks match them.
What is missing
In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:
getSpanName()(lines 511–522) switches onprovider + ":" + lastPathSegment. ForPOST /v1/messages/batches, the last segment is"batches", which doesn't match thePROVIDER_NAME_ANTHROPIC + ":messages"case, so it falls through todefault -> lastSegment, yielding the span name"batches"instead of something like"anthropic.messages.batches.create".tagAnthropicRequest()(lines 208–255) unconditionally setsspan_attributes = {"type":"llm"}(line 216) even though a batch-create call isn't itself a model invocation. It only populatesmetadata.modelandbraintrust.input_jsonwhenrequestJson.has("model")/requestJson.has("messages")(lines 231, 235) — but aBatchCreateParamsrequest body has neither at the top level; it has arequests[]array where each entry is{"custom_id": "...", "params": {"model": ..., "messages": ..., "max_tokens": ..., ...}}. Every one of the (potentially many) inline generation requests in the batch is silently dropped from the span.tagAnthropicResponse()(lines 258–301) dumps the whole response body asoutput_json(harmless for batch-create, since the response is just batch job metadata:id,processing_status,request_counts,results_url) and only extractsmetricsfrom a top-levelusageobject (line 270), which a batch-create/retrieve response never has — so no metrics are produced (correctly, since none exist yet at creation time, but there's also no instrumentation of the batch results retrieval, which is where the actual per-requestusageandmessageoutputs become available).- There is no wrapping at all of
batches().results(...)/resultsStreaming(...), so even after a batch completes, iterating its per-request results (each containing acustom_idand either a succeededmessagewith full usage, or an error) produces zero spans — unlike a normalmessages().create()call, none of the individual generation results in a batch get any Braintrust span. - No test or example anywhere in the repo exercises
client.messages().batches()in any form (confirmed via repo-wide grep forbatch/Batchinanthropic_2_2_0— zero matches).
Braintrust docs status: supported (in a sibling SDK) / not_found (for Java)
- The Java section of https://www.braintrust.dev/docs/integrations/ai-providers/anthropic states only: "Braintrust emits spans for the Anthropic Messages API. Each span captures the input messages and response content," with a spans table listing exactly one row — "Anthropic Messages API spans" covering
messages().create()calls including streaming. Batches are not mentioned for Java at all: not_found. - The Python section of the same page states: "Braintrust emits spans for the Anthropic SDK's messages, batches, and managed agents APIs," with a spans-table row for
anthropic.messages.batches.*explicitly covering the Batches API: supported (in the Python SDK). This establishes Message Batches tracing as an existing, documented Braintrust capability that the Java SDK has not brought to parity.
Upstream sources
- Anthropic Java SDK Batches sub-client:
client.messages().batches().create(BatchCreateParams),.retrieve(...),.list(...),.cancel(...),.resultsStreaming(BatchResultsParams)— official Java code samples in the batch-processing guide: https://platform.claude.com/docs/en/build-with-claude/batch-processing (mirrored at https://docs.anthropic.com/en/docs/build-with-claude/batch-processing) - Batch create request shape — each
requests[]entry hascustom_id+paramscontaining "the standard Messages API parameters" (model, max_tokens, messages, system, tools, thinking, etc.), i.e. inline generation requests, not file/S3 references: same guide, "Prepare and create your batch" section - Batch results shape — streamed JSONL keyed by
custom_id, each withresult.typeofsucceeded(containing the fullmessage),errored,canceled, orexpired: https://platform.claude.com/docs/en/api/messages/batches/results - Currency/stability — Anthropic's official API release notes confirm this is an actively maintained, GA surface (e.g. March 30, 2026 entry raising
max_tokensto 300k specifically "on the Message Batches API"; the Claude-on-AWS platform announcement listing "the full Messages API, Files API, Message Batches API, Claude Managed Agents" as core stable surfaces): https://platform.claude.com/docs/en/release-notes/api
Local files inspected
braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java— lines 208–255 (tagAnthropicRequest: no top-levelmodel/messagesin a batch-create body, so metadata/input_json stay empty), lines 258–301 (tagAnthropicResponse: whole-body dump plususage-only metrics, batch-create/retrieve responses have neitherusagenor useful whole-body content), lines 511–522 (getSpanName:"batches"path segment falls through to the genericdefaultcase)braintrust-sdk/instrumentation/anthropic_2_2_0/src/main/java/dev/braintrust/instrumentation/anthropic/v2_2_0/TracingHttpClient.javaandContextCapturingProxy.java— generic transport-swap/service-graph-following wrapper; produces a span for any Anthropic HTTP call including batches, but has no batch-specific logicbraintrust-sdk/instrumentation/anthropic_2_2_0/src/test/java/dev/braintrust/instrumentation/anthropic/v2_2_0/BraintrustAnthropicTest.javaandBraintrustAnthropicPromptCachingTest.java— no test exercisesbatches()in any formexamples/anthropic-instrumentation/— no batch example exists- Repo-wide grep for
batch/Batchunderbraintrust-sdk/instrumentation/anthropic_2_2_0/— zero matches
- Ngôn ngữ chính
- Java
- Star
- 21
- Fork
- 5
- Merge trung bình
- 2 ngày 7 giờ
- Pull request đã merge (30 ngày)
- 8
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của braintrustdata/braintrust-sdk-java
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Tất cả issue của braintrustdata/braintrust-sdk-java
Issue tương tự
-
certification
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 80/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
[BUG] ECR GetAuthorizationToken returns a proxyEndpoint for the default region, not the request's Đang mởbug ecr
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
Needs: Triage Type: Feature request
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
AntennaPod/AntennaPod#8794 ·
-
agentic-workflows
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
github/copilot-sdk#2760 ·