[Live API] gemini-3.8-live-extended-thinking: announces tool call then says "I apologize, but a system error occurred." with no functionCall frame - 61% of tool runs (n=64)
I maintainer di solito rispondono entro 1 giorno
@Venkaiahbabuneelam ci sta già lavorando.
Dal 28/9/2026.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
Summary
On models/gemini-3.8-live-extended-thinking, tool calls fail server-side in a way that is undetectable from the client: the model first speaks a normal turn committing to the call (e.g. "Let me update that file…"), then 1–3 seconds later says "I apologize, but a system error occurred." — and no functionCall frame is ever sent. The turn closes with a normal turnComplete and no APIError/close frame, so exception-based detection is impossible.
We have measured this over n=64 tool-scenario runs: 25 pass / 39 fail = 61% failure rate, across all thinking levels. A fully independent raw-WebSocket reproduction (no application code, minimal client) reports the identical signature: https://github.com/masudl-hub/theoremai/issues/25
Environment
google-genai==2.25.0, Python 3.14, Live API over websocket- Model
models/gemini-3.8-live-extended-thinking,thinking_leveltried athigh,low,medium - Tools declared as
functionDeclarationswithbehavior: NON_BLOCKING(13 tools in our production session; failures reproduce with a single trivial tool as well) - Session also uses
session_resumption,context_window_compression,media_resolution: MEDIUM— but the minimal reproduction in the thread above uses none of these and still fails
Reproduction steps
- Open a Live session on
gemini-3.8-live-extended-thinkingwith one tool declared (behavior: NON_BLOCKING). - Send a text/audio prompt that requires the tool (in our harness: "edit the file … replace the word done with finished").
- Observed: audio/content turn announces the action, then the apology text;
toolCall/functionCallframe never arrives;turnCompletefires normally; no client-side error object. - Expected: a
functionCallframe so the client can dispatch and respond.
Repeat several times — the failure is intermittent within the same session config (some runs fully clean, others fail 3–4 times in a row).
Measured rates (our E2E harness, one row per run)
| Config | Pass / Fail |
|---|---|
extended + tool, thinking_level=high |
16 / 29 |
extended + tool, thinking_level=low |
4 / 4 |
extended + tool, thinking_level=medium |
1 / 2 |
| extended + tool, unlabelled legacy runs | 4 / 4 |
| extended + tool, total | 25 / 39 = 61% fail (n=64) |
| extended, chat only (no tools) | 2 / 0 |
base gemini-3.8-live + same tools |
2 / 2 (small-n; earlier A/B: 0/3 errors vs 2/3 for extended) |
A fresh 10-run batch on 2026-09-28 was worse (3/10 pass), confirming high run-to-run variance rather than a fixed rate.
Additional contract observations
behavior: BLOCKINGis rejected at session setup with close code 1007: "BLOCKING function calls are not supported for this model." (reported in the thread above), soNON_BLOCKINGis the only path for this model.- Because there is no error signal, the only practical detection is text-signature based (the apology phrase). We mitigated by injecting a short text nudge after the model goes IDLE ("verify the actual result, then re-call the tool"): this recovered 4 of 6 detected failures in one batch, but it is a workaround, not a fix, and it cannot help clients that don't watch the transcript.
- This looks related to #2827 ("model composes a function call in thinking but never emits it"), possibly the same underlying class surfaced differently (apology text vs. fabricated answer).
Ask
- Investigate why the extended-thinking Live model aborts tool-call emission server-side while the base
gemini-3.8-livemodel on the same setup succeeds more often. - Consider whether the failure should surface as an explicit error/event instead of plain assistant text, so clients can retry or fail over.
Happy to share redacted session logs (full websocket transcripts + per-run timing).
- Lingua principale
- Python
- Stelle
- 4k
- Fork
- 1k
- Merge medio
- 1g 18h
- PR unite (30g)
- 52
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di googleapis/python-genai
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
googleapis/python-genai#3051 ·
I maintainer di solito rispondono entro 1 giorno
-
[Bug]: Unsubscripted typing.List and typing.Dict crash convert_if_exist_pydantic_model in AFCForse già presa @chauvuusvn l’ha presa 2 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
googleapis/python-genai#3044 ·
I maintainer di solito rispondono entro 1 giorno
-
Name collision in google.genai.interactions: triggers.Interaction shadows response model Interaction in static type checkersForse già presa @Venkaiahbabuneelam l’ha presa 7 giorni fa. Apertapriority: p2 type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
googleapis/python-genai#3013 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 78/100
googleapis/python-genai#3031 · 2 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
Live `AsyncSession.send_tool_response` fails with `TypeError` when `FunctionResponse.parts` carries inline bytesForse già presa @Venkaiahbabuneelam l’ha presa 5 giorni fa. Apertapriority: p2 type: bug
Difficoltà 3/5 1-2 giorni Idoneità per principianti 72/100
googleapis/python-genai#3022 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di googleapis/python-genai
Issue simili
-
needs triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
json_params_matcher fails on falsy top-level JSON primitives (0, False, "")Forse già presa @mayureshsonawane17 l’ha presa oggi. ApertaWaiting for: Product Owner
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 5 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
I maintainer di solito rispondono entro 1 giorno
-
Add .devin pluginAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
ayghri/i-have-adhd#249 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
modelscope/FunASR#3762 ·
I maintainer di solito rispondono entro 1 giorno