Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Gemini API latency has recently increased significantly, especially with store=true

未关闭
#3,059 2 条评论 0 个 reaction 已指派 1 人 在 GitHub 查看

维护者通常 1 天内回复

@Venkaiahbabuneelam 已经在做这个了。

开始于 2026年10月6日。

评估

这个 Issue 还没有评估数据。

描述

priority: p3 status:awaiting user response type: question

We have noticed a significant increase in the time before Gemini responses start, particularly when conversation storage is enabled (store=true).

For Gemini 3.5 Flash-Lite, the median time to first text increased from approximately 1.2 seconds to 4.0 seconds. This means approximately 2.8 seconds of additional waiting before any text appears. We also found this storage-related delay in other Gemini models.

Historical latency comparison

All values below are medians in seconds.

Model Earlier measurement October 6, 2026 Increase
gemini-3.5-flash-lite: time to first text 1.196 (September 7) 4.010 +2.814
gemini-3.5-flash-lite: answer ready 1.440 (September 7) 4.246 +2.806
gemini-3.7-flash: time to first text 1.255 (September 3) 5.135 +3.880

The Flash-Lite figures cover 10 requests per measurement; the 3.7 Flash figures cover 5. All of these requests used store=true.

Storage-related delay across models

On October 6, we observed the following median times to full API completion:

Model store=true store=false Additional time with storage
gemini-3.5-flash-lite 3.822 1.010 +2.812
gemini-3.5-flash 4.852 1.854 +2.998
gemini-3.7-flash 7.258 2.289 +4.969
gemini-3.8-flash 4.822 1.735 +3.087

Each entry covers 5 requests, with all 40 requests succeeding. These are full completion times, whereas the historical table above measures time to first text and answer readiness. The storage-related delay appears across all four models; the historical increase differs by model.

Even a minimal hello prompt shows the delay

Using Gemini 3.5 Flash-Lite with only the prompt 只說hello (say hello only):

Metric store=true store=false Difference
Median time to first text 3.340 0.625 +2.715
Median time to full API completion 4.993 0.683 +4.310

All 20 requests succeeded, with 10 requests per setting. Even this very short prompt showed approximately 2.7 seconds of additional waiting before the first text appeared.

We observed no rate-limit errors, application retries, or fallback models in these batches. Direct curl requests also reproduced the storage-related delay.

The main concern is the increased waiting before the response starts, which noticeably affects interactive use. We cannot determine the internal cause from client-side measurements.

Has there been a recent change or a known latency issue with the Interactions API when store=true? Is the additional delay expected, and is there a way to reduce it while keeping conversation storage enabled?

主要语言
Python
星标
4k
派生
1k
平均合并
1 天 17 小时
30 天内合并 PR
54

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

googleapis/python-genai 的其他 Issue

查看 googleapis/python-genai 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。