Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Gemini API latency has recently increased significantly, especially with store=true

Aperta
#3,059 2 commenti 0 reazioni 1 assegnatario Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

@Venkaiahbabuneelam ci sta già lavorando.

Dal 6/10/2026.

Valutazione

Questa issue non è ancora stata valutata.

Descrizione

priority: p3 status:awaiting user response type: question

We have noticed a significant increase in the time before Gemini responses start, particularly when conversation storage is enabled (store=true).

For Gemini 3.5 Flash-Lite, the median time to first text increased from approximately 1.2 seconds to 4.0 seconds. This means approximately 2.8 seconds of additional waiting before any text appears. We also found this storage-related delay in other Gemini models.

Historical latency comparison

All values below are medians in seconds.

Model Earlier measurement October 6, 2026 Increase
gemini-3.5-flash-lite: time to first text 1.196 (September 7) 4.010 +2.814
gemini-3.5-flash-lite: answer ready 1.440 (September 7) 4.246 +2.806
gemini-3.7-flash: time to first text 1.255 (September 3) 5.135 +3.880

The Flash-Lite figures cover 10 requests per measurement; the 3.7 Flash figures cover 5. All of these requests used store=true.

Storage-related delay across models

On October 6, we observed the following median times to full API completion:

Model store=true store=false Additional time with storage
gemini-3.5-flash-lite 3.822 1.010 +2.812
gemini-3.5-flash 4.852 1.854 +2.998
gemini-3.7-flash 7.258 2.289 +4.969
gemini-3.8-flash 4.822 1.735 +3.087

Each entry covers 5 requests, with all 40 requests succeeding. These are full completion times, whereas the historical table above measures time to first text and answer readiness. The storage-related delay appears across all four models; the historical increase differs by model.

Even a minimal hello prompt shows the delay

Using Gemini 3.5 Flash-Lite with only the prompt 只說hello (say hello only):

Metric store=true store=false Difference
Median time to first text 3.340 0.625 +2.715
Median time to full API completion 4.993 0.683 +4.310

All 20 requests succeeded, with 10 requests per setting. Even this very short prompt showed approximately 2.7 seconds of additional waiting before the first text appeared.

We observed no rate-limit errors, application retries, or fallback models in these batches. Direct curl requests also reproduced the storage-related delay.

The main concern is the increased waiting before the response starts, which noticeably affects interactive use. We cannot determine the internal cause from client-side measurements.

Has there been a recent change or a known latency issue with the Interactions API when store=true? Is the additional delay expected, and is there a way to reduce it while keeping conversation storage enabled?

Lingua principale
Python
Stelle
4k
Fork
1k
Merge medio
1g 18h
PR unite (30g)
65

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di googleapis/python-genai

Tutte le issue di googleapis/python-genai

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.