Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[bot] Google GenAI: streaming responses drop `url_context_metadata` that non-streaming responses preserve

未关闭 适合新手
#774 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
86/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
python
领域
backend, testing

调研方向

从 py/src/braintrust/integrations/google_genai/tracing.py 中的 _aggregate_generate_content_chunks() 开始,然后将其 candidate 字段处理与 py/src/braintrust/integrations/google_genai/test_google_genai.py 中现有的 grounding 元数据测试进行比较。为同步和异步路径中的流式 url_context 元数据添加覆盖,并验证元数据存在于生成的 span 输出中。

由索引模型根据 Issue 内容生成。

描述

<!-- provider-gap-audit: google-genai-streaming-url-context-metadata -->

Summary

When the Gemini url_context tool is used with generate_content_stream() / agenerate_content_stream(), the per-URL retrieval metadata (candidate.url_context_metadata) that Google's SDK returns is silently dropped from the Braintrust span output. The equivalent metadata for the google_search tool (candidate.grounding_metadata) is preserved in the exact same code path — so this is a fidelity gap between two structurally-analogous tool result types within the same function, not a "feature never built" gap.

Non-streaming generate_content() calls do not have this problem: the raw GenerateContentResponse object (including url_context_metadata) is logged as-is, so nothing is lost there.

What is missing

_aggregate_generate_content_chunks() in py/src/braintrust/integrations/google_genai/tracing.py (used by both the sync and async streaming wrappers) manually reconstructs a candidate_dict from the accumulated chunks, copying over only an explicit allowlist of candidate fields:

candidate_dict = {"content": {"parts": parts, "role": "model"}}

if hasattr(candidate, "finish_reason"):
    candidate_dict["finish_reason"] = candidate.finish_reason
if hasattr(candidate, "safety_ratings"):
    candidate_dict["safety_ratings"] = candidate.safety_ratings
if hasattr(candidate, "grounding_metadata") and candidate.grounding_metadata:
    candidate_dict["grounding_metadata"] = candidate.grounding_metadata

(py/src/braintrust/integrations/google_genai/tracing.py:695-709)

candidate.url_context_metadata — the field Google's own docs say to inspect to see "which URLs the model retrieved" when the url_context tool is enabled — is never copied into candidate_dict, so it never reaches the logged span output for streaming calls. Any user who calls client.models.generate_content_stream(..., config=GenerateContentConfig(tools=[{"url_context": {}}])) gets a span with no record of which URLs were actually fetched, even though the same call via generate_content() (non-streaming) would show it.

This is the same class of field (candidate.<x>_metadata describing what a built-in tool did) as grounding_metadata, which is explicitly captured here and has dedicated test coverage (test_google_search_grounding / test_google_search_grounding_async in test_google_genai.py). There is no equivalent test for url_context, and a full-file grep for url_context or code_execution in test_google_genai.py returns zero matches — confirming there is no regression coverage that would have caught this gap.

Note: the separate _TOOL_CALL_TYPES/_TOOL_RESULT_TYPES constants and interaction-tool-span logic elsewhere in the same file (tracing.py:54-69, :877-969) do already generically recognize url_context_call/url_context_result and code_execution_call/code_execution_result — that mechanism belongs to the newer content-item/"interactions" API surface and is unrelated to the classic generate_content_stream() candidate-based aggregation described above, which is the specific path where the metadata is lost.

Braintrust docs status

not_found — https://www.braintrust.dev/docs/integrations/ai-providers/google-genai (and the general https://www.braintrust.dev/docs/guides/tracing) do not document url_context tool support or grounding/citation-style metadata capture at all, streaming or otherwise.

Upstream sources

Local repo files inspected

  • py/src/braintrust/integrations/google_genai/tracing.py:
    • _aggregate_generate_content_chunks() (~lines 641-722) — builds candidate_dict for streaming span output; copies finish_reason, safety_ratings, grounding_metadata but not url_context_metadata
    • _gc_process_result() (~lines 573-581) — non-streaming path; returns the raw GenerateContentResponse, so no loss there
    • _TOOL_CALL_TYPES / _TOOL_RESULT_TYPES (~lines 54-69) and the interaction-tool-span logic (~lines 877-969) — confirmed this is a separate code path (content-item/interactions API) unrelated to the candidate-based streaming aggregation gap above
  • py/src/braintrust/integrations/google_genai/test_google_genai.py:
    • test_google_search_grounding / test_google_search_grounding_async (~lines 1451, 1557) and _assert_grounding_metadata (~line 1411) — dedicated grounding-metadata test exists for google_search only
    • Full-file grep for url_context and code_execution — zero matches, confirming no test coverage for either tool type
主要语言
Python
星标
19
派生
17
平均合并
1 天 5 小时
30 天内合并 PR
61

贡献指南

打开贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

braintrustdata/braintrust-sdk-python 的其他 Issue

查看 braintrustdata/braintrust-sdk-python 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。