Hacktoberfest 2026 : les issues que les mainteneurs ont marquées pour octobre, ouvertes et accessibles aux débutants. Parcourir les issues Hacktoberfest

after_model_callback: a replacement LlmResponse drops `usage_metadata`, erasing the model call from token accounting

Fermée Adaptée aux débutants
#7,451 0 commentaires 0 réactions 0 personnes assignées Voir sur GitHub

Les mainteneurs répondent en général sous 5 jours

@zhuhongd y travaille déjà.

Depuis le 7/10/2026.

  • #7449 par @zhuhongd — ouverte

Évaluation

Difficulté
2/5
Temps estimé
1-3 heures
Accessibilité débutants
68/100
Type d'issue
Bug
Clarté
Clairement spécifiée
Activité
Active
Stack technique
python
Domaine
backend

Piste de recherche

Commencez par flows/llm_flows/core/_finalizer.py, où, selon l’issue, partial et turn_complete sont déjà hérités et grounding_metadata est pris en charge. Vérifiez les tests du finalizer concernant les réponses de remplacement créées par callback, puis ajoutez des tests pour un usage_metadata non défini et pour un usage_metadata fourni explicitement. Le travail est terminé lorsque le usage hérité apparaît dans les événements finaux et persistés, tandis qu’une valeur de remplacement explicite est conservée.

Rédigé par le modèle d'indexation à partir du texte de l'issue.

Description

🔴 Required Information

Describe the Bug:

Follow-up to #7035. That fix (#7036, 39e4538) makes a callback-built
replacement inherit the streaming-control fields partial and
turn_complete; usage_metadata was deliberately left out of its scope.

The same mechanism still drops it. When an after_model_callback returns a
rebuilt LlmResponse (the documented contract) and does not set
usage_metadata, the replacement carries None. finalize_model_response_event
copies only non-None fields into the Event, so the yielded final event and
the event persisted to the session have no usage for that model call.
Everything that reads usage from events loses it: the BigQuery analytics
plugin's event path, the A2A converters, agent_test_runner, and any client
doing per-call cost attribution from session history. Plugins that read
llm_response.usage_metadata inside their own after_model_callback still
see it, because plugin callbacks run before the agent callback, which is why
the loss is easy to miss in logs.

usage_metadata measures the model call (prompt, candidate and cached token
counts billed by the provider), not the content of the response. Replacing
the content does not change what the call cost, so there is no reading under
which None is the more accurate value. It applies in streaming and
non-streaming mode alike.

To Reproduce:

Self-contained, no API key (fake model reports usage on its final response,
as Gemini does):

pip install google-adk==2.11.0 && python repro_usage_loss.py
# repro_usage_loss.py
"""Repro: after_model_callback replacement drops usage_metadata on current main."""
import asyncio
from typing import AsyncGenerator

from google.adk.agents import LlmAgent
from google.adk.agents.run_config import RunConfig, StreamingMode
from google.adk.models.base_llm import BaseLlm
from google.adk.models.llm_request import LlmRequest
from google.adk.models.llm_response import LlmResponse
from google.adk.runners import InMemoryRunner
from google.genai import types

USAGE = types.GenerateContentResponseUsageMetadata(
    prompt_token_count=120, candidates_token_count=8, total_token_count=128)


def _resp(text, partial=None, usage=None):
  return LlmResponse(
      content=types.Content(role="model", parts=[types.Part(text=text)]),
      partial=partial, usage_metadata=usage)


class FakeLlm(BaseLlm):
  @classmethod
  def supported_models(cls): return [".*"]
  async def generate_content_async(self, llm_request: LlmRequest, stream=False
                                   ) -> AsyncGenerator[LlmResponse, None]:
    if stream:
      for d in ["Hello ", "world."]:
        yield _resp(d, partial=True)
    yield _resp("Hello world.", usage=USAGE)  # provider reports usage on final


def rebuild_cb(callback_context, llm_response):
  if not (llm_response.content and llm_response.content.parts): return None
  t = llm_response.content.parts[0].text or ""
  return _resp(t.replace("world", "[REDACTED]"))


async def run(label, cb, stream):
  agent = LlmAgent(name="a", model=FakeLlm(model="fake"), after_model_callback=cb)
  runner = InMemoryRunner(agent=agent, app_name="r")
  s = await runner.session_service.create_session(app_name="r", user_id="u")
  cfg = RunConfig(streaming_mode=StreamingMode.SSE if stream else StreamingMode.NONE)
  finals = []
  async for ev in runner.run_async(user_id="u", session_id=s.id,
      new_message=types.Content(role="user", parts=[types.Part(text="hi")]), run_config=cfg):
    if not ev.partial: finals.append(ev)
  stored = await runner.session_service.get_session(app_name="r", user_id="u", session_id=s.id)
  persisted = [e for e in stored.events if e.author == "a"]
  tok = lambda e: e.usage_metadata.total_token_count if e.usage_metadata else None
  print(f"{label:22} final_event.usage={tok(finals[-1])!s:5} persisted.usage={tok(persisted[-1])!s:5} partial_flags={[e.partial for e in persisted]}")

async def main():
  import google.adk; print("google-adk", google.adk.__version__)
  await run("CONTROL non-stream", None, False)
  await run("REBUILD non-stream", rebuild_cb, False)
  await run("CONTROL SSE", None, True)
  await run("REBUILD SSE", rebuild_cb, True)
asyncio.run(main())

Actual output on google-adk 2.11.0 (latest release; also reproduced on main @ 42a17a9f):

google-adk 2.11.0
CONTROL non-stream     final_event.usage=128   persisted.usage=128   partial_flags=[None]
REBUILD non-stream     final_event.usage=None  persisted.usage=None  partial_flags=[None]
CONTROL SSE            final_event.usage=128   persisted.usage=128   partial_flags=[None]
REBUILD SSE            final_event.usage=None  persisted.usage=None  partial_flags=[None]

(The partial flags are correct in all four runs, confirming #7036 landed
and that this is the remaining gap.)

Expected behavior:

A replacement that leaves usage_metadata unset inherits it from the
response it replaces, so the persisted event still reports 128 tokens. A
replacement that sets its own usage_metadata (e.g. a callback that made an
additional model call and wants to report combined usage) is respected.

Suggested fix:

Extend the inheritance already applied in flows/llm_flows/core/_finalizer.py for
partial/turn_complete to usage_metadata, same copy-on-inherit
semantics (no mutation of the callback-owned object, explicit values win).
Scope deliberately limited to this one field: grounding_metadata already
has its own handling in the same function, and finish_reason/error_code
are left to the callback because a guardrail replacement may legitimately
change finish semantics. PR: #7449

Desktop:

  • OS: macOS (Darwin 27.0)
  • Python: 3.13.12
  • google-adk: 2.11.0 (release) and main @ 42a17a9f
Langage dominant
Python
Étoiles
21.8k
Forks
4.1k
Merge moyen
1 j 13 h
PR mergées (30 j)
6

Préparer son environnement

Par où commencer

  1. Lisez l'issue en entier, puis le guide de contribution du projet.
  2. Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
  3. Forkez le dépôt et travaillez sur une branche.
  4. Ouvrez une pull request qui référence le numéro de l'issue.

Autres issues de google/adk-python

Toutes les issues de google/adk-python

Issues similaires

Plus d'issues Python

Recevez les nouvelles issues par e-mail

Un résumé court des issues GitHub adaptées aux débutants.