Gemini API latency has recently increased significantly, especially with store=true
I maintainer di solito rispondono entro 1 giorno
@Venkaiahbabuneelam ci sta già lavorando.
Dal 6/10/2026.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
We have noticed a significant increase in the time before Gemini responses start, particularly when conversation storage is enabled (store=true).
For Gemini 3.5 Flash-Lite, the median time to first text increased from approximately 1.2 seconds to 4.0 seconds. This means approximately 2.8 seconds of additional waiting before any text appears. We also found this storage-related delay in other Gemini models.
Historical latency comparison
All values below are medians in seconds.
| Model | Earlier measurement | October 6, 2026 | Increase |
|---|---|---|---|
| gemini-3.5-flash-lite: time to first text | 1.196 (September 7) | 4.010 | +2.814 |
| gemini-3.5-flash-lite: answer ready | 1.440 (September 7) | 4.246 | +2.806 |
| gemini-3.7-flash: time to first text | 1.255 (September 3) | 5.135 | +3.880 |
The Flash-Lite figures cover 10 requests per measurement; the 3.7 Flash figures cover 5. All of these requests used store=true.
Storage-related delay across models
On October 6, we observed the following median times to full API completion:
| Model | store=true | store=false | Additional time with storage |
|---|---|---|---|
| gemini-3.5-flash-lite | 3.822 | 1.010 | +2.812 |
| gemini-3.5-flash | 4.852 | 1.854 | +2.998 |
| gemini-3.7-flash | 7.258 | 2.289 | +4.969 |
| gemini-3.8-flash | 4.822 | 1.735 | +3.087 |
Each entry covers 5 requests, with all 40 requests succeeding. These are full completion times, whereas the historical table above measures time to first text and answer readiness. The storage-related delay appears across all four models; the historical increase differs by model.
Even a minimal hello prompt shows the delay
Using Gemini 3.5 Flash-Lite with only the prompt 只說hello (say hello only):
| Metric | store=true | store=false | Difference |
|---|---|---|---|
| Median time to first text | 3.340 | 0.625 | +2.715 |
| Median time to full API completion | 4.993 | 0.683 | +4.310 |
All 20 requests succeeded, with 10 requests per setting. Even this very short prompt showed approximately 2.7 seconds of additional waiting before the first text appeared.
We observed no rate-limit errors, application retries, or fallback models in these batches. Direct curl requests also reproduced the storage-related delay.
The main concern is the increased waiting before the response starts, which noticeably affects interactive use. We cannot determine the internal cause from client-side measurements.
Has there been a recent change or a known latency issue with the Interactions API when store=true? Is the additional delay expected, and is there a way to reduce it while keeping conversation storage enabled?
- Lingua principale
- Python
- Stelle
- 4k
- Fork
- 1k
- Merge medio
- 1g 18h
- PR unite (30g)
- 65
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di googleapis/python-genai
-
Curated history keeps half a user turn when the model turn is invalidForse già presa @Venkaiahbabuneelam l’ha presa 4 giorni fa. Apertapriority: p2 status:awaiting user response type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
googleapis/python-genai#3051 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
[Bug]: Unsubscripted typing.List and typing.Dict crash convert_if_exist_pydantic_model in AFCForse già presa @Venkaiahbabuneelam l’ha presa 4 giorni fa. Apertapriority: p2 type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
googleapis/python-genai#3044 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Name collision in google.genai.interactions: triggers.Interaction shadows response model Interaction in static type checkersForse già presa @Venkaiahbabuneelam l’ha presa 11 giorni fa. Apertapriority: p2 type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
googleapis/python-genai#3013 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
googleapis/python-genai#3078 ·
I maintainer di solito rispondono entro 1 giorno
-
Video understanding on the Interactions API: files registered from GCS (files.register_files) produce inflated, fabricated event lists; the same bytes uploaded (files.upload) do notForse già presa @Venkaiahbabuneelam l’ha presa 1 giorno fa. Apertapriority: p2 type: bug
googleapis/python-genai#3072 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di googleapis/python-genai
Issue simili
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0Apertadocumentation priority: low size: S
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
I maintainer di solito rispondono entro 1 giorno
-
bug help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
I maintainer di solito rispondono entro 1 giorno
-
Broken link in index.rstApertadocumentation
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 65/100
ansys/pydpf-core#3547 ·
I maintainer di solito rispondono entro 1 giorno
-
core
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
vectorize-io/hindsight#5457 ·
I maintainer di solito rispondono entro 1 giorno
-
[Bug]: LangChain drops OpenAI Responses text blocks from session recordingForse già presa @ktz03 l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
volcengine/OpenViking#5806 ·
I maintainer di solito rispondono entro 1 giorno