[bot] Anthropic Message Batches API is not instrumented (mis-tagged as a generic LLM span)
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- java
- Domain
- api, observability
Research direction
Start with InstrumentationSemConv.java, especially tagAnthropicRequest(), tagAnthropicResponse(), and getSpanName(), then inspect TracingHttpClient.java and ContextCapturingProxy.java to understand the existing span flow. Review BraintrustAnthropicTest.java and the Anthropic batch request and results shapes in the linked upstream documentation. Done means batch operations have meaningful spans and metadata, and resultsStreaming() coverage verifies successful and errored per-request results.
Written by the indexing model from the issue text.
Description
Summary
The Anthropic instrumentation module (anthropic_2_2_0) only has shape-aware tagging for the Messages API (messages().create/streaming). It has no support at all for the Message Batches API (client.messages().batches()), a stable, GA, actively-developed part of the Anthropic Java SDK for submitting many message-generation requests as a single asynchronous job.
Because the generic transport-swap (TracingHttpClient) still intercepts the batch HTTP calls, a batches().create(...) call does produce a span — but it is actively mis-tagged rather than simply absent: it gets span_attributes.type = "llm" (as if it were a real model call), a low-information span name ("batches", from the path-segment fallback), no model in metadata, and no input_json at all — because the request/response shapes for batches are structurally different from a single Messages call and none of the existing field checks match them.
What is missing
In braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java:
getSpanName()(lines 511–522) switches onprovider + ":" + lastPathSegment. ForPOST /v1/messages/batches, the last segment is"batches", which doesn't match thePROVIDER_NAME_ANTHROPIC + ":messages"case, so it falls through todefault -> lastSegment, yielding the span name"batches"instead of something like"anthropic.messages.batches.create".tagAnthropicRequest()(lines 208–255) unconditionally setsspan_attributes = {"type":"llm"}(line 216) even though a batch-create call isn't itself a model invocation. It only populatesmetadata.modelandbraintrust.input_jsonwhenrequestJson.has("model")/requestJson.has("messages")(lines 231, 235) — but aBatchCreateParamsrequest body has neither at the top level; it has arequests[]array where each entry is{"custom_id": "...", "params": {"model": ..., "messages": ..., "max_tokens": ..., ...}}. Every one of the (potentially many) inline generation requests in the batch is silently dropped from the span.tagAnthropicResponse()(lines 258–301) dumps the whole response body asoutput_json(harmless for batch-create, since the response is just batch job metadata:id,processing_status,request_counts,results_url) and only extractsmetricsfrom a top-levelusageobject (line 270), which a batch-create/retrieve response never has — so no metrics are produced (correctly, since none exist yet at creation time, but there's also no instrumentation of the batch results retrieval, which is where the actual per-requestusageandmessageoutputs become available).- There is no wrapping at all of
batches().results(...)/resultsStreaming(...), so even after a batch completes, iterating its per-request results (each containing acustom_idand either a succeededmessagewith full usage, or an error) produces zero spans — unlike a normalmessages().create()call, none of the individual generation results in a batch get any Braintrust span. - No test or example anywhere in the repo exercises
client.messages().batches()in any form (confirmed via repo-wide grep forbatch/Batchinanthropic_2_2_0— zero matches).
Braintrust docs status: supported (in a sibling SDK) / not_found (for Java)
- The Java section of https://www.braintrust.dev/docs/integrations/ai-providers/anthropic states only: "Braintrust emits spans for the Anthropic Messages API. Each span captures the input messages and response content," with a spans table listing exactly one row — "Anthropic Messages API spans" covering
messages().create()calls including streaming. Batches are not mentioned for Java at all: not_found. - The Python section of the same page states: "Braintrust emits spans for the Anthropic SDK's messages, batches, and managed agents APIs," with a spans-table row for
anthropic.messages.batches.*explicitly covering the Batches API: supported (in the Python SDK). This establishes Message Batches tracing as an existing, documented Braintrust capability that the Java SDK has not brought to parity.
Upstream sources
- Anthropic Java SDK Batches sub-client:
client.messages().batches().create(BatchCreateParams),.retrieve(...),.list(...),.cancel(...),.resultsStreaming(BatchResultsParams)— official Java code samples in the batch-processing guide: https://platform.claude.com/docs/en/build-with-claude/batch-processing (mirrored at https://docs.anthropic.com/en/docs/build-with-claude/batch-processing) - Batch create request shape — each
requests[]entry hascustom_id+paramscontaining "the standard Messages API parameters" (model, max_tokens, messages, system, tools, thinking, etc.), i.e. inline generation requests, not file/S3 references: same guide, "Prepare and create your batch" section - Batch results shape — streamed JSONL keyed by
custom_id, each withresult.typeofsucceeded(containing the fullmessage),errored,canceled, orexpired: https://platform.claude.com/docs/en/api/messages/batches/results - Currency/stability — Anthropic's official API release notes confirm this is an actively maintained, GA surface (e.g. March 30, 2026 entry raising
max_tokensto 300k specifically "on the Message Batches API"; the Claude-on-AWS platform announcement listing "the full Messages API, Files API, Message Batches API, Claude Managed Agents" as core stable surfaces): https://platform.claude.com/docs/en/release-notes/api
Local files inspected
braintrust-sdk/src/main/java/dev/braintrust/instrumentation/InstrumentationSemConv.java— lines 208–255 (tagAnthropicRequest: no top-levelmodel/messagesin a batch-create body, so metadata/input_json stay empty), lines 258–301 (tagAnthropicResponse: whole-body dump plususage-only metrics, batch-create/retrieve responses have neitherusagenor useful whole-body content), lines 511–522 (getSpanName:"batches"path segment falls through to the genericdefaultcase)braintrust-sdk/instrumentation/anthropic_2_2_0/src/main/java/dev/braintrust/instrumentation/anthropic/v2_2_0/TracingHttpClient.javaandContextCapturingProxy.java— generic transport-swap/service-graph-following wrapper; produces a span for any Anthropic HTTP call including batches, but has no batch-specific logicbraintrust-sdk/instrumentation/anthropic_2_2_0/src/test/java/dev/braintrust/instrumentation/anthropic/v2_2_0/BraintrustAnthropicTest.javaandBraintrustAnthropicPromptCachingTest.java— no test exercisesbatches()in any formexamples/anthropic-instrumentation/— no batch example exists- Repo-wide grep for
batch/Batchunderbraintrust-sdk/instrumentation/anthropic_2_2_0/— zero matches
- Dominant language
- Java
- Stars
- 21
- Forks
- 5
- Avg merge
- 2d 7h
- Merged PRs (30d)
- 8
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from braintrustdata/braintrust-sdk-java
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
All issues in braintrustdata/braintrust-sdk-java
Similar issues
-
certification
Difficulty 1/5 Under an hour Newbie friendliness 80/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
[BUG] ECR GetAuthorizationToken returns a proxyEndpoint for the default region, not the request's Openbug ecr
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Needs: Triage Type: Feature request
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
AntennaPod/AntennaPod#8794 ·
-
agentic-workflows
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
github/copilot-sdk#2760 ·