Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Gemini API latency has recently increased significantly, especially with store=true

Đang mở
#3,059 2 bình luận 0 reaction 1 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

@Venkaiahbabuneelam đang làm issue này rồi.

Từ ngày 6/10/2026.

Đánh giá

Issue này chưa được đánh giá.

Mô tả

priority: p3 status:awaiting user response type: question

We have noticed a significant increase in the time before Gemini responses start, particularly when conversation storage is enabled (store=true).

For Gemini 3.5 Flash-Lite, the median time to first text increased from approximately 1.2 seconds to 4.0 seconds. This means approximately 2.8 seconds of additional waiting before any text appears. We also found this storage-related delay in other Gemini models.

Historical latency comparison

All values below are medians in seconds.

Model Earlier measurement October 6, 2026 Increase
gemini-3.5-flash-lite: time to first text 1.196 (September 7) 4.010 +2.814
gemini-3.5-flash-lite: answer ready 1.440 (September 7) 4.246 +2.806
gemini-3.7-flash: time to first text 1.255 (September 3) 5.135 +3.880

The Flash-Lite figures cover 10 requests per measurement; the 3.7 Flash figures cover 5. All of these requests used store=true.

Storage-related delay across models

On October 6, we observed the following median times to full API completion:

Model store=true store=false Additional time with storage
gemini-3.5-flash-lite 3.822 1.010 +2.812
gemini-3.5-flash 4.852 1.854 +2.998
gemini-3.7-flash 7.258 2.289 +4.969
gemini-3.8-flash 4.822 1.735 +3.087

Each entry covers 5 requests, with all 40 requests succeeding. These are full completion times, whereas the historical table above measures time to first text and answer readiness. The storage-related delay appears across all four models; the historical increase differs by model.

Even a minimal hello prompt shows the delay

Using Gemini 3.5 Flash-Lite with only the prompt 只說hello (say hello only):

Metric store=true store=false Difference
Median time to first text 3.340 0.625 +2.715
Median time to full API completion 4.993 0.683 +4.310

All 20 requests succeeded, with 10 requests per setting. Even this very short prompt showed approximately 2.7 seconds of additional waiting before the first text appeared.

We observed no rate-limit errors, application retries, or fallback models in these batches. Direct curl requests also reproduced the storage-related delay.

The main concern is the increased waiting before the response starts, which noticeably affects interactive use. We cannot determine the internal cause from client-side measurements.

Has there been a recent change or a known latency issue with the Interactions API when store=true? Is the additional delay expected, and is there a way to reduce it while keeping conversation storage enabled?

Ngôn ngữ chính
Python
Star
4k
Fork
1k
Merge trung bình
1 ngày 17 giờ
Pull request đã merge (30 ngày)
54

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của googleapis/python-genai

Tất cả issue của googleapis/python-genai

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.