Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

LiteLLM integration does not report cached, reasoning, or cache-write token usage

Aperta
#5,455 4 commenti 0 reazioni 1 assegnatario Vedi su GitHub

@alexander-alderman-webb ci sta già lavorando.

Dal 23/2/2026.

Valutazione

Questa issue non è ancora stata valutata.

Descrizione

Bug Python Spans
How do you use Sentry?

Sentry Saas (sentry.io)

Version

2.52.0

Steps to Reproduce
  1. Initialize Sentry with the LiteLLM integration and tracing enabled
  2. Make a completion call through LiteLLM to a provider that supports prompt caching (e.g., OpenAI, Anthropic, etc.)
  3. Inspect the resulting span data in Sentry's AI Agents dashboard
Expected Result

The span should include all available token usage detail attributes, just like the OpenAI and Anthropic integrations do:

  • gen_ai.usage.input_tokens (total input tokens)
  • gen_ai.usage.input_tokens.cached (cached input tokens, subset of total)
  • gen_ai.usage.input_tokens.cache_write (cache write tokens, if available)
  • gen_ai.usage.output_tokens (total output tokens)
  • gen_ai.usage.output_tokens.reasoning (reasoning tokens, subset of total)
  • gen_ai.usage.total_tokens

This data is necessary for Sentry to correctly calculate model costs using the formula documented here:

input cost = (input_tokens - cached_tokens) x input_rate + cached_tokens x cached_rate

Without cached/reasoning token breakdown, all tokens are charged at the full standard rate, producing inaccurate cost estimates.

Actual Result

The LiteLLM integration's _success_callback only extracts three basic fields:

record_token_usage(
    span,
    input_tokens=getattr(usage, "prompt_tokens", None),
    output_tokens=getattr(usage, "completion_tokens", None),
    total_tokens=getattr(usage, "total_tokens", None),
)

The input_tokens_cached, input_tokens_cache_write, and output_tokens_reasoning parameters of record_token_usage() are never passed. Therefore, cost calculations in the AI Agents dashboard overestimate costs for cached-heavy workloads (all input tokens billed at the full rate) and misattribute output vs. reasoning token costs.

Lingua principale
Python
Stelle
2.2k
Fork
672
Merge medio
22h 47m
PR unite (30g)
224

Guida per i contributori

Apri la guida per i contributori

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di getsentry/sentry-python

Tutte le issue di getsentry/sentry-python

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.