Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

otelMiddleware: media generation spans never capture prompt, input media or output, even with captureContent

Closed
#1,526 0 comments 0 reactions 1 assignee View on GitHub

Maintainers usually reply within 1 day

@AlemTuzlak is already working on this.

Since Sep 27, 2026.

Assessment

This issue has not been assessed yet.

Description

waiting-on: maintainer

Summary

otelMiddleware opens one CLIENT span for each media activity call (generateImage, generateVideo, generateAudio, generateSpeech, generateTranscription, …), added via #720. That span carries the provider, the model, the operation name and usage. It never records what was asked for or what came back, even with captureContent: true:

  • no prompt
  • no input media: reference images, start and end frames, source audio for transcription
  • no output: generated image, video or audio URLs, or the transcript text

In PostHog, Langfuse or Datadog every media generation therefore shows up as an empty Input/Output pair with a cost attached. For image-to-video or reference-to-video, where the result depends mostly on the input images, there is nothing in the trace to debug from.

Root cause

startMediaSpan / endMediaSpan (packages/ai/src/middlewares/otel.ts:343–384) only set gen_ai.system, gen_ai.operation.name, gen_ai.request.model and usage. captureContent is only checked on the chat paths.

The inputs are already on the context: GenerationMiddlewareContext.artifactInputs holds the activity inputs. The media path never reads it, and the result is never inspected in the terminal hook.

The only workaround is attributeEnricher / onSpanEnd in every app, reimplementing the per-activity input and output shapes that the library already knows.

Proposal

When captureContent is on, the media span writes the same attributes the chat iteration spans do:

  • gen_ai.input.messages / langfuse.observation.input: the prompt as a text part, plus each input media reference as a URI part ({ type: 'uri', modality, uri }), reusing whatever part serialisation #1525 lands on.
  • gen_ai.output.messages / langfuse.observation.output: output media URLs as URI parts, and transcript text as a text part.
  • Inline or base64 inputs and outputs follow the same rule as #1525: a placeholder by default, and an opt-in hook to swap in a URL. redact and maxContentLength still apply.

Related

  • #720: added the media span (closed)
  • #1525: captureContent drops non-text parts on chat spans

Version

@tanstack/ai 0.58.0 (current main has the same code).

I'm happy to open a PR.

Dominant language
TypeScript
Stars
3.1k
Forks
340
Avg merge
2d 10h
Merged PRs (30d)
174

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from TanStack/ai

All issues in TanStack/ai

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.