Gemini API latency has recently increased significantly, especially with store=true
维护者通常 1 天内回复
@Venkaiahbabuneelam 已经在做这个了。
开始于 2026年10月6日。
评估
这个 Issue 还没有评估数据。
描述
We have noticed a significant increase in the time before Gemini responses start, particularly when conversation storage is enabled (store=true).
For Gemini 3.5 Flash-Lite, the median time to first text increased from approximately 1.2 seconds to 4.0 seconds. This means approximately 2.8 seconds of additional waiting before any text appears. We also found this storage-related delay in other Gemini models.
Historical latency comparison
All values below are medians in seconds.
| Model | Earlier measurement | October 6, 2026 | Increase |
|---|---|---|---|
| gemini-3.5-flash-lite: time to first text | 1.196 (September 7) | 4.010 | +2.814 |
| gemini-3.5-flash-lite: answer ready | 1.440 (September 7) | 4.246 | +2.806 |
| gemini-3.7-flash: time to first text | 1.255 (September 3) | 5.135 | +3.880 |
The Flash-Lite figures cover 10 requests per measurement; the 3.7 Flash figures cover 5. All of these requests used store=true.
Storage-related delay across models
On October 6, we observed the following median times to full API completion:
| Model | store=true | store=false | Additional time with storage |
|---|---|---|---|
| gemini-3.5-flash-lite | 3.822 | 1.010 | +2.812 |
| gemini-3.5-flash | 4.852 | 1.854 | +2.998 |
| gemini-3.7-flash | 7.258 | 2.289 | +4.969 |
| gemini-3.8-flash | 4.822 | 1.735 | +3.087 |
Each entry covers 5 requests, with all 40 requests succeeding. These are full completion times, whereas the historical table above measures time to first text and answer readiness. The storage-related delay appears across all four models; the historical increase differs by model.
Even a minimal hello prompt shows the delay
Using Gemini 3.5 Flash-Lite with only the prompt 只說hello (say hello only):
| Metric | store=true | store=false | Difference |
|---|---|---|---|
| Median time to first text | 3.340 | 0.625 | +2.715 |
| Median time to full API completion | 4.993 | 0.683 | +4.310 |
All 20 requests succeeded, with 10 requests per setting. Even this very short prompt showed approximately 2.7 seconds of additional waiting before the first text appeared.
We observed no rate-limit errors, application retries, or fallback models in these batches. Direct curl requests also reproduced the storage-related delay.
The main concern is the increased waiting before the response starts, which noticeably affects interactive use. We cannot determine the internal cause from client-side measurements.
Has there been a recent change or a known latency issue with the Interactions API when store=true? Is the additional delay expected, and is there a way to reduce it while keeping conversation storage enabled?
- 主要语言
- Python
- 星标
- 4k
- 派生
- 1k
- 平均合并
- 1 天 17 小时
- 30 天内合并 PR
- 54
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
googleapis/python-genai 的其他 Issue
-
Curated history keeps half a user turn when the model turn is invalid可能已有人在做 @Venkaiahbabuneelam 于 2 天前认领。 未关闭priority: p2 status:awaiting user response type: bug
难度 2/5 1-3 小时 新手友好度 62/100
googleapis/python-genai#3051 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
[Bug]: Unsubscripted typing.List and typing.Dict crash convert_if_exist_pydantic_model in AFC可能已有人在做 @chauvuusvn 于 4 天前认领。 未关闭priority: p2 type: bug
难度 2/5 1-3 小时 新手友好度 82/100
googleapis/python-genai#3044 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
Name collision in google.genai.interactions: triggers.Interaction shadows response model Interaction in static type checkers可能已有人在做 @Venkaiahbabuneelam 于 9 天前认领。 未关闭priority: p2 type: bug
难度 2/5 1-3 小时 新手友好度 78/100
googleapis/python-genai#3013 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
Callable tools: optional scalar parameters (`int | None`) get `"type": "object"` in `parameters_json_schema`, so Gemini sends `{}`可能已有人在做 @Venkaiahbabuneelam 于 1 天前认领。 未关闭priority: p2 type: bug
googleapis/python-genai#3060 · 1 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
-
live.connect() (Vertex AI) resolves ADC and refreshes the OAuth token synchronously on the event loop on every connect可能已有人在做 @Venkaiahbabuneelam 于 1 天前认领。 未关闭priority: p2 type: bug
难度 3/5 1-2 天 新手友好度 74/100
googleapis/python-genai#3056 · 2 条评论 · 已指派 1 人 ·
维护者通常 1 天内回复
查看 googleapis/python-genai 的全部 Issue
相似的 Issue
-
area/install reliability
难度 2/5 1-3 小时 新手友好度 75/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 83/100
FluidNumerics/fluid-walk-blocker#191 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 62/100
TransformerLensOrg/TransformerLens#1868 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 85/100
climate-analytics-lab/jax-gcm#1057 ·
维护者通常 1 天内回复