after_model_callback: a replacement LlmResponse drops `usage_metadata`, erasing the model call from token accounting
Maintainer thường phản hồi trong vòng 6 ngày
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 68/100
Hướng nghiên cứu
Bắt đầu với flows/llm_flows/core/_finalizer.py, nơi issue cho biết partial và turn_complete đã được kế thừa, còn grounding_metadata được xử lý. Kiểm tra các bài kiểm thử finalizer cho phản hồi thay thế được tạo bằng callback, rồi bổ sung kiểm thử cho cả trường hợp usage_metadata chưa được đặt và được cung cấp tường minh. Hoàn tất khi usage được kế thừa xuất hiện trong các sự kiện cuối cùng và được lưu trữ, còn giá trị thay thế tường minh được giữ nguyên.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
🔴 Required Information
Describe the Bug:
Follow-up to #7035. That fix (#7036, 39e4538) makes a callback-built
replacement inherit the streaming-control fields partial and
turn_complete; usage_metadata was deliberately left out of its scope.
The same mechanism still drops it. When an after_model_callback returns a
rebuilt LlmResponse (the documented contract) and does not set
usage_metadata, the replacement carries None. finalize_model_response_event
copies only non-None fields into the Event, so the yielded final event and
the event persisted to the session have no usage for that model call.
Everything that reads usage from events loses it: the BigQuery analytics
plugin's event path, the A2A converters, agent_test_runner, and any client
doing per-call cost attribution from session history. Plugins that read
llm_response.usage_metadata inside their own after_model_callback still
see it, because plugin callbacks run before the agent callback, which is why
the loss is easy to miss in logs.
usage_metadata measures the model call (prompt, candidate and cached token
counts billed by the provider), not the content of the response. Replacing
the content does not change what the call cost, so there is no reading under
which None is the more accurate value. It applies in streaming and
non-streaming mode alike.
To Reproduce:
Self-contained, no API key (fake model reports usage on its final response,
as Gemini does):
pip install google-adk==2.11.0 && python repro_usage_loss.py
# repro_usage_loss.py
"""Repro: after_model_callback replacement drops usage_metadata on current main."""
import asyncio
from typing import AsyncGenerator
from google.adk.agents import LlmAgent
from google.adk.agents.run_config import RunConfig, StreamingMode
from google.adk.models.base_llm import BaseLlm
from google.adk.models.llm_request import LlmRequest
from google.adk.models.llm_response import LlmResponse
from google.adk.runners import InMemoryRunner
from google.genai import types
USAGE = types.GenerateContentResponseUsageMetadata(
prompt_token_count=120, candidates_token_count=8, total_token_count=128)
def _resp(text, partial=None, usage=None):
return LlmResponse(
content=types.Content(role="model", parts=[types.Part(text=text)]),
partial=partial, usage_metadata=usage)
class FakeLlm(BaseLlm):
@classmethod
def supported_models(cls): return [".*"]
async def generate_content_async(self, llm_request: LlmRequest, stream=False
) -> AsyncGenerator[LlmResponse, None]:
if stream:
for d in ["Hello ", "world."]:
yield _resp(d, partial=True)
yield _resp("Hello world.", usage=USAGE) # provider reports usage on final
def rebuild_cb(callback_context, llm_response):
if not (llm_response.content and llm_response.content.parts): return None
t = llm_response.content.parts[0].text or ""
return _resp(t.replace("world", "[REDACTED]"))
async def run(label, cb, stream):
agent = LlmAgent(name="a", model=FakeLlm(model="fake"), after_model_callback=cb)
runner = InMemoryRunner(agent=agent, app_name="r")
s = await runner.session_service.create_session(app_name="r", user_id="u")
cfg = RunConfig(streaming_mode=StreamingMode.SSE if stream else StreamingMode.NONE)
finals = []
async for ev in runner.run_async(user_id="u", session_id=s.id,
new_message=types.Content(role="user", parts=[types.Part(text="hi")]), run_config=cfg):
if not ev.partial: finals.append(ev)
stored = await runner.session_service.get_session(app_name="r", user_id="u", session_id=s.id)
persisted = [e for e in stored.events if e.author == "a"]
tok = lambda e: e.usage_metadata.total_token_count if e.usage_metadata else None
print(f"{label:22} final_event.usage={tok(finals[-1])!s:5} persisted.usage={tok(persisted[-1])!s:5} partial_flags={[e.partial for e in persisted]}")
async def main():
import google.adk; print("google-adk", google.adk.__version__)
await run("CONTROL non-stream", None, False)
await run("REBUILD non-stream", rebuild_cb, False)
await run("CONTROL SSE", None, True)
await run("REBUILD SSE", rebuild_cb, True)
asyncio.run(main())
Actual output on google-adk 2.11.0 (latest release; also reproduced on main @ 42a17a9f):
google-adk 2.11.0
CONTROL non-stream final_event.usage=128 persisted.usage=128 partial_flags=[None]
REBUILD non-stream final_event.usage=None persisted.usage=None partial_flags=[None]
CONTROL SSE final_event.usage=128 persisted.usage=128 partial_flags=[None]
REBUILD SSE final_event.usage=None persisted.usage=None partial_flags=[None]
(The partial flags are correct in all four runs, confirming #7036 landed
and that this is the remaining gap.)
Expected behavior:
A replacement that leaves usage_metadata unset inherits it from the
response it replaces, so the persisted event still reports 128 tokens. A
replacement that sets its own usage_metadata (e.g. a callback that made an
additional model call and wants to report combined usage) is respected.
Suggested fix:
Extend the inheritance already applied in flows/llm_flows/core/_finalizer.py for
partial/turn_complete to usage_metadata, same copy-on-inherit
semantics (no mutation of the callback-owned object, explicit values win).
Scope deliberately limited to this one field: grounding_metadata already
has its own handling in the same function, and finish_reason/error_code
are left to the callback because a guardrail replacement may legitimately
change finish semantics. PR: #7449
Desktop:
- OS: macOS (Darwin 27.0)
- Python: 3.13.12
- google-adk: 2.11.0 (release) and main @ 42a17a9f
- Ngôn ngữ chính
- Python
- Star
- 21.6k
- Fork
- 4k
- Merge trung bình
- 1 ngày 13 giờ
- Pull request đã merge (30 ngày)
- 6
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của google/adk-python
-
GKE code executor unit tests fail with kubernetes 37.0.0, turning main CI redCó thể đã có người làm @vetler đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
google/adk-python#7443 ·
Maintainer thường phản hồi trong vòng 6 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
google/adk-python#7433 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 6 ngày
-
CI Mypy Check flags an existing streaming_utils.py error as new because the PR run reuses the baseline's mypy cacheCó thể đã có người làm @DeanChensj đã nhận 2 ngày trước. Đang mởneeds review
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
google/adk-python#7409 · 1 bình luận · 2 người được giao ·
Maintainer thường phản hồi trong vòng 6 ngày
-
A2aAgentExecutor sends the raw exception text to the A2A caller when the run failsCó thể đã có người làm @sanketpatil06 đã nhận 3 ngày trước. Đang mởa2a request clarification
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
google/adk-python#7385 · 2 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 6 ngày
-
Please support mermaid 12 (inbuild elk) in `adk web`Có thể đã có người làm @sanketpatil06 đã nhận 3 ngày trước. Đang mởneeds review web
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
google/adk-python#7381 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 6 ngày
Tất cả issue của google/adk-python
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
Maintainer thường phản hồi trong vòng 1 ngày
-
SR_SECURITY_DESCRIPTOR.fromString drops the SACL when no DACL is presentCó thể đã có người làm @paul7436 đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
equinor/fmu-sumo-uploader#302 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
modelscope/evalscope#1821 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Sanity on ansible-core devel fails: ignore-2.23.txt references the removed import-3.9 testCó thể đã có người làm @yurnov đã nhận hôm nay. Đang mởneeds_triage
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 91/100
ansible-collections/kubernetes.core#1275 ·
Maintainer thường phản hồi trong vòng 1 ngày