[Live API] gemini-3.8-live-extended-thinking: announces tool call then says "I apologize, but a system error occurred." with no functionCall frame - 61% of tool runs (n=64)
维护者通常 1 天内回复
@Venkaiahbabuneelam 已经在做这个了。
开始于 2026年9月28日。
评估
这个 Issue 还没有评估数据。
描述
Summary
On models/gemini-3.8-live-extended-thinking, tool calls fail server-side in a way that is undetectable from the client: the model first speaks a normal turn committing to the call (e.g. "Let me update that file…"), then 1–3 seconds later says "I apologize, but a system error occurred." — and no functionCall frame is ever sent. The turn closes with a normal turnComplete and no APIError/close frame, so exception-based detection is impossible.
We have measured this over n=64 tool-scenario runs: 25 pass / 39 fail = 61% failure rate, across all thinking levels. A fully independent raw-WebSocket reproduction (no application code, minimal client) reports the identical signature: https://github.com/masudl-hub/theoremai/issues/25
Environment
google-genai==2.25.0, Python 3.14, Live API over websocket- Model
models/gemini-3.8-live-extended-thinking,thinking_leveltried athigh,low,medium - Tools declared as
functionDeclarationswithbehavior: NON_BLOCKING(13 tools in our production session; failures reproduce with a single trivial tool as well) - Session also uses
session_resumption,context_window_compression,media_resolution: MEDIUM— but the minimal reproduction in the thread above uses none of these and still fails
Reproduction steps
- Open a Live session on
gemini-3.8-live-extended-thinkingwith one tool declared (behavior: NON_BLOCKING). - Send a text/audio prompt that requires the tool (in our harness: "edit the file … replace the word done with finished").
- Observed: audio/content turn announces the action, then the apology text;
toolCall/functionCallframe never arrives;turnCompletefires normally; no client-side error object. - Expected: a
functionCallframe so the client can dispatch and respond.
Repeat several times — the failure is intermittent within the same session config (some runs fully clean, others fail 3–4 times in a row).
Measured rates (our E2E harness, one row per run)
| Config | Pass / Fail |
|---|---|
extended + tool, thinking_level=high |
16 / 29 |
extended + tool, thinking_level=low |
4 / 4 |
extended + tool, thinking_level=medium |
1 / 2 |
| extended + tool, unlabelled legacy runs | 4 / 4 |
| extended + tool, total | 25 / 39 = 61% fail (n=64) |
| extended, chat only (no tools) | 2 / 0 |
base gemini-3.8-live + same tools |
2 / 2 (small-n; earlier A/B: 0/3 errors vs 2/3 for extended) |
A fresh 10-run batch on 2026-09-28 was worse (3/10 pass), confirming high run-to-run variance rather than a fixed rate.
Additional contract observations
behavior: BLOCKINGis rejected at session setup with close code 1007: "BLOCKING function calls are not supported for this model." (reported in the thread above), soNON_BLOCKINGis the only path for this model.- Because there is no error signal, the only practical detection is text-signature based (the apology phrase). We mitigated by injecting a short text nudge after the model goes IDLE ("verify the actual result, then re-call the tool"): this recovered 4 of 6 detected failures in one batch, but it is a workaround, not a fix, and it cannot help clients that don't watch the transcript.
- This looks related to #2827 ("model composes a function call in thinking but never emits it"), possibly the same underlying class surfaced differently (apology text vs. fabricated answer).
Ask
- Investigate why the extended-thinking Live model aborts tool-call emission server-side while the base
gemini-3.8-livemodel on the same setup succeeds more often. - Consider whether the failure should surface as an explicit error/event instead of plain assistant text, so clients can retry or fail over.
Happy to share redacted session logs (full websocket transcripts + per-run timing).
- 主要语言
- Python
- 星标
- 4k
- 派生
- 1k
- 平均合并
- 1 天 18 小时
- 30 天内合并 PR
- 65
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
googleapis/python-genai 的其他 Issue
-
Curated history keeps half a user turn when the model turn is invalid可能已有人在做 @Venkaiahbabuneelam 于 4 天前认领。 未关闭priority: p2 status:awaiting user response type: bug
难度 2/5 1-3 小时 新手友好度 62/100
googleapis/python-genai#3051 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[Bug]: Unsubscripted typing.List and typing.Dict crash convert_if_exist_pydantic_model in AFC可能已有人在做 @Venkaiahbabuneelam 于 4 天前认领。 未关闭priority: p2 type: bug
难度 2/5 1-3 小时 新手友好度 82/100
googleapis/python-genai#3044 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
Name collision in google.genai.interactions: triggers.Interaction shadows response model Interaction in static type checkers可能已有人在做 @Venkaiahbabuneelam 于 11 天前认领。 未关闭priority: p2 type: bug
难度 2/5 1-3 小时 新手友好度 78/100
googleapis/python-genai#3013 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
难度 4/5 3-5 天 新手友好度 55/100
googleapis/python-genai#3078 ·
维护者通常 1 天内回复
-
Video understanding on the Interactions API: files registered from GCS (files.register_files) produce inflated, fabricated event lists; the same bytes uploaded (files.upload) do not可能已有人在做 @Venkaiahbabuneelam 于 1 天前认领。 未关闭priority: p2 type: bug
googleapis/python-genai#3072 · 已指派 1 人 ·
维护者通常 1 天内回复
查看 googleapis/python-genai 的全部 Issue
相似的 Issue
-
难度 1/5 1 小时以内 新手友好度 85/100
MystenLabs/MemWal#1163 · 2 条评论 ·
维护者通常 1 天内回复
-
infertopics leaves new nodes without a topic when untopiced neighbours outnumber topiced ones可能已有人在做 @moneebullah25 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 70/100
FinanceFlash/unvibecode#218 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 75/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 70/100
NVIDIA/earth2studio#1241 ·
维护者通常 3 天内回复