[Live API] gemini-3.8-live-extended-thinking: announces tool call then says "I apologize, but a system error occurred." with no functionCall frame - 61% of tool runs (n=64)
Los mantenedores suelen responder en 1 día
@Venkaiahbabuneelam ya está trabajando en esto.
Desde el 28/9/2026.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Summary
On models/gemini-3.8-live-extended-thinking, tool calls fail server-side in a way that is undetectable from the client: the model first speaks a normal turn committing to the call (e.g. "Let me update that file…"), then 1–3 seconds later says "I apologize, but a system error occurred." — and no functionCall frame is ever sent. The turn closes with a normal turnComplete and no APIError/close frame, so exception-based detection is impossible.
We have measured this over n=64 tool-scenario runs: 25 pass / 39 fail = 61% failure rate, across all thinking levels. A fully independent raw-WebSocket reproduction (no application code, minimal client) reports the identical signature: https://github.com/masudl-hub/theoremai/issues/25
Environment
google-genai==2.25.0, Python 3.14, Live API over websocket- Model
models/gemini-3.8-live-extended-thinking,thinking_leveltried athigh,low,medium - Tools declared as
functionDeclarationswithbehavior: NON_BLOCKING(13 tools in our production session; failures reproduce with a single trivial tool as well) - Session also uses
session_resumption,context_window_compression,media_resolution: MEDIUM— but the minimal reproduction in the thread above uses none of these and still fails
Reproduction steps
- Open a Live session on
gemini-3.8-live-extended-thinkingwith one tool declared (behavior: NON_BLOCKING). - Send a text/audio prompt that requires the tool (in our harness: "edit the file … replace the word done with finished").
- Observed: audio/content turn announces the action, then the apology text;
toolCall/functionCallframe never arrives;turnCompletefires normally; no client-side error object. - Expected: a
functionCallframe so the client can dispatch and respond.
Repeat several times — the failure is intermittent within the same session config (some runs fully clean, others fail 3–4 times in a row).
Measured rates (our E2E harness, one row per run)
| Config | Pass / Fail |
|---|---|
extended + tool, thinking_level=high |
16 / 29 |
extended + tool, thinking_level=low |
4 / 4 |
extended + tool, thinking_level=medium |
1 / 2 |
| extended + tool, unlabelled legacy runs | 4 / 4 |
| extended + tool, total | 25 / 39 = 61% fail (n=64) |
| extended, chat only (no tools) | 2 / 0 |
base gemini-3.8-live + same tools |
2 / 2 (small-n; earlier A/B: 0/3 errors vs 2/3 for extended) |
A fresh 10-run batch on 2026-09-28 was worse (3/10 pass), confirming high run-to-run variance rather than a fixed rate.
Additional contract observations
behavior: BLOCKINGis rejected at session setup with close code 1007: "BLOCKING function calls are not supported for this model." (reported in the thread above), soNON_BLOCKINGis the only path for this model.- Because there is no error signal, the only practical detection is text-signature based (the apology phrase). We mitigated by injecting a short text nudge after the model goes IDLE ("verify the actual result, then re-call the tool"): this recovered 4 of 6 detected failures in one batch, but it is a workaround, not a fix, and it cannot help clients that don't watch the transcript.
- This looks related to #2827 ("model composes a function call in thinking but never emits it"), possibly the same underlying class surfaced differently (apology text vs. fabricated answer).
Ask
- Investigate why the extended-thinking Live model aborts tool-call emission server-side while the base
gemini-3.8-livemodel on the same setup succeeds more often. - Consider whether the failure should surface as an explicit error/event instead of plain assistant text, so clients can retry or fail over.
Happy to share redacted session logs (full websocket transcripts + per-run timing).
- Lenguaje dominante
- Python
- Estrellas
- 4k
- Forks
- 1k
- Merge medio
- 1 d 17 h
- PR fusionados (30 d)
- 52
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de googleapis/python-genai
-
priority: p2 status:awaiting user response type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
googleapis/python-genai#3051 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
[Bug]: Unsubscripted typing.List and typing.Dict crash convert_if_exist_pydantic_model in AFCPosiblemente ocupada @chauvuusvn la tomó hace 3 días. Abiertopriority: p2 type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
googleapis/python-genai#3044 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
Name collision in google.genai.interactions: triggers.Interaction shadows response model Interaction in static type checkersPosiblemente ocupada @Venkaiahbabuneelam la tomó hace 8 días. Abiertopriority: p2 type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
googleapis/python-genai#3013 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
Dificultad 3/5 1-2 días Aptitud para principiantes 74/100
googleapis/python-genai#3056 ·
Los mantenedores suelen responder en 1 día
-
Vertex batch embeddings (gemini-embedding-001): no way to attach billing labels; job and row labels never reach Cloud BillingPosiblemente ocupada @Venkaiahbabuneelam la tomó hace 1 día. Abiertopriority: p3 type: question
googleapis/python-genai#3053 · 1 asignado ·
Los mantenedores suelen responder en 1 día
Todos los issues de googleapis/python-genai
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Juniper/ansible-junos-stdlib#904 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
pollen-robotics/reachy_mini#1457 ·
Los mantenedores suelen responder en 1 día
-
area:runtime good first issue
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
WATonomous/wato_f1tenth#39 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
FireDynamics/fdsreader#123 ·
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100