Gemini API latency has recently increased significantly, especially with store=true
Maintainer thường phản hồi trong vòng 1 ngày
@Venkaiahbabuneelam đang làm issue này rồi.
Từ ngày 6/10/2026.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
We have noticed a significant increase in the time before Gemini responses start, particularly when conversation storage is enabled (store=true).
For Gemini 3.5 Flash-Lite, the median time to first text increased from approximately 1.2 seconds to 4.0 seconds. This means approximately 2.8 seconds of additional waiting before any text appears. We also found this storage-related delay in other Gemini models.
Historical latency comparison
All values below are medians in seconds.
| Model | Earlier measurement | October 6, 2026 | Increase |
|---|---|---|---|
| gemini-3.5-flash-lite: time to first text | 1.196 (September 7) | 4.010 | +2.814 |
| gemini-3.5-flash-lite: answer ready | 1.440 (September 7) | 4.246 | +2.806 |
| gemini-3.7-flash: time to first text | 1.255 (September 3) | 5.135 | +3.880 |
The Flash-Lite figures cover 10 requests per measurement; the 3.7 Flash figures cover 5. All of these requests used store=true.
Storage-related delay across models
On October 6, we observed the following median times to full API completion:
| Model | store=true | store=false | Additional time with storage |
|---|---|---|---|
| gemini-3.5-flash-lite | 3.822 | 1.010 | +2.812 |
| gemini-3.5-flash | 4.852 | 1.854 | +2.998 |
| gemini-3.7-flash | 7.258 | 2.289 | +4.969 |
| gemini-3.8-flash | 4.822 | 1.735 | +3.087 |
Each entry covers 5 requests, with all 40 requests succeeding. These are full completion times, whereas the historical table above measures time to first text and answer readiness. The storage-related delay appears across all four models; the historical increase differs by model.
Even a minimal hello prompt shows the delay
Using Gemini 3.5 Flash-Lite with only the prompt 只說hello (say hello only):
| Metric | store=true | store=false | Difference |
|---|---|---|---|
| Median time to first text | 3.340 | 0.625 | +2.715 |
| Median time to full API completion | 4.993 | 0.683 | +4.310 |
All 20 requests succeeded, with 10 requests per setting. Even this very short prompt showed approximately 2.7 seconds of additional waiting before the first text appeared.
We observed no rate-limit errors, application retries, or fallback models in these batches. Direct curl requests also reproduced the storage-related delay.
The main concern is the increased waiting before the response starts, which noticeably affects interactive use. We cannot determine the internal cause from client-side measurements.
Has there been a recent change or a known latency issue with the Interactions API when store=true? Is the additional delay expected, and is there a way to reduce it while keeping conversation storage enabled?
- Ngôn ngữ chính
- Python
- Star
- 4k
- Fork
- 1k
- Merge trung bình
- 1 ngày 17 giờ
- Pull request đã merge (30 ngày)
- 54
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của googleapis/python-genai
-
Curated history keeps half a user turn when the model turn is invalidCó thể đã có người làm @Venkaiahbabuneelam đã nhận 2 ngày trước. Đang mởpriority: p2 status:awaiting user response type: bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
googleapis/python-genai#3051 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[Bug]: Unsubscripted typing.List and typing.Dict crash convert_if_exist_pydantic_model in AFCCó thể đã có người làm @chauvuusvn đã nhận 4 ngày trước. Đang mởpriority: p2 type: bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
googleapis/python-genai#3044 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Name collision in google.genai.interactions: triggers.Interaction shadows response model Interaction in static type checkersCó thể đã có người làm @Venkaiahbabuneelam đã nhận 9 ngày trước. Đang mởpriority: p2 type: bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
googleapis/python-genai#3013 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Callable tools: optional scalar parameters (`int | None`) get `"type": "object"` in `parameters_json_schema`, so Gemini sends `{}`Có thể đã có người làm @Venkaiahbabuneelam đã nhận 1 ngày trước. Đang mởpriority: p2 type: bug
googleapis/python-genai#3060 · 1 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
-
live.connect() (Vertex AI) resolves ADC and refreshes the OAuth token synchronously on the event loop on every connectCó thể đã có người làm @Venkaiahbabuneelam đã nhận 1 ngày trước. Đang mởpriority: p2 type: bug
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 74/100
googleapis/python-genai#3056 · 2 bình luận · 1 người được giao ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của googleapis/python-genai
Issue tương tự
-
changelog investigate
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
ramnes/notion-sdk-py#409 ·
-
good first issue help wanted
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
lindicaphxag-tech/kaggle#28 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 62/100
BSData/horus-heresy-3rd-edition#3211 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
Maintainer thường phản hồi trong vòng 1 ngày
-
bug tests
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày