Video understanding on the Interactions API: files registered from GCS (files.register_files) produce inflated, fabricated event lists; the same bytes uploaded (files.upload) do not
I maintainer di solito rispondono entro 1 giorno
@Venkaiahbabuneelam ci sta già lavorando.
Dal 8/10/2026.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
I use video understanding with gemini-3.1-pro-preview to count how many times a specific chain of events occurs in a video. The video is sampled at 4 fps (static video processing, processing: {"type": "static", "fps": 4}) rather than the default 1 fps, because the events are short. The model returns one structured JSON entry per occurrence, with start and end timestamps and a short description. Our test video has 13 counts of the chain (hand-counted).
This application had been running without issue on the previous API (models.generate_content), with videos both registered from GCS and uploaded through the File API. The problem appeared when I migrated the same requests to the Interactions API.
On the Interactions API, the answer depends on how the video is set as input:
- Registered from Cloud Storage (
client.files.register_files(uris=["gs://…"])): 25 of 60 calls (about 40%) returned an inflated, fabricated list. These answers report 24–52 occurrences instead of 13. - Uploaded with File API directly (
client.files.upload(file=…)): 0 of 20 calls did this. The mean count was 13.6–13.7. - Old API (models.generate_content) with the same model, prompt, schema and fps: 0 of 30 testing calls did this, whether the file was registered or uploaded. (Neither did I encounter this issue with non-testing data)
Everything else about the requests in those tests was identical.
What an "inflated, fabricated list" looks like
These answers aren't small counting errors. The list stops reflecting the video:
- Far too many entries: 24 to 52 occurrences reported where there are 13, roughly 2–4× the true count.
Timestamps on a fixed grid: entries are spaced evenly every 2–3 seconds from the start of the video to the end, whether or not anything happens at those times. - Copy-pasted descriptions: most entries share one identical, generic description (for example "Consistent performance." repeated 41 times out of 51), instead of describing what actually happens in each occurrence.
- The count is decided up front: the thought summary typically states a large total early on (for example "parsed into 52 distinct repetitions") and then fills the list to match it.
Correct answers, by contrast, have unevenly spaced timestamps that follow the real events, and a distinct description for each entry.
For counting, we classify an answer as inflated when it reports 24 or more occurrences, close to twice the true 13. Correct answers never came near that: across all clean arms the counts ranged from 8 to 19.
Environment details
- Both macOS 27 & Cloud Run container
- SDK:
google-genai==2.28.0(Interactions), andgoogle-genai==2.11.0for thegenerate_contentbaseline - Python: 3.11
- Model:
gemini-3.1-pro-preview - Test Video: 126 s, 1280×720, 30 fps, MP4
- Request settings:
- static video processing at
fps: 4 - no
resolutionset (the API default) service_tier="standard"store=False- default thinking, with
thinking_summaries="auto" - structured output via
response_format(type: text,mime_type: application/json, JSON Schema from a Pydantic model)
- static video processing at
- Dates: 2026-10-07 and 2026-10-08
Steps to reproduce
import json, os
from google import genai
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
# A) registered from GCS
reg = client.files.register_files(auth=gcs_credentials, uris=["gs://<bucket>/videos/clip.mp4"]).files[0]
# B) uploaded (same bytes)
up = client.files.upload(file="clip.mp4", config={"mime_type": "video/mp4"})
# (wait until up.state == ACTIVE)
def count_events(file_uri):
ix = client.interactions.create(
model="gemini-3.1-pro-preview",
input=[
{"type": "video", "uri": file_uri, "mime_type": "video/mp4",
"processing": {"type": "static", "fps": 4.0}},
{"type": "text", "text": PROMPT}, # "identify every occurrence of <chain of events>…"
],
response_format={"type": "text", "mime_type": "application/json",
"schema": EventsAnalysis.model_json_schema()},
generation_config={"thinking_summaries": "auto"},
service_tier="standard",
store=False,
)
return len(json.loads(ix.output_text)["events"])
Ruled out by changing one variable at a time and running set of 10 calls:
- video codec (HEVC vs H.264)
- audio processing (loudness-normalised vs untouched)
- keyframe spacing (8–42 s vs 1 s)
- JSON Schema vs SDK-converted schema
- explicit
temperature
Other observations about registered files
- No video metadata. After
register_filesthe file isACTIVEimmediately andfiles.get(...).video_metadataisNone. Afterfiles.uploadthe file goesPROCESSING → ACTIVE(about 45 s for 25 MB) and reportsvideo_metadata={'videoDuration': '126s'}. - Same token count. Both deliveries bill the same input tokens (39,801 vs 39,802), so the model receives the same amount of video. What differs seems to be the frames or timestamps it receives, not how many. That would fit the evenly spaced, ungrounded timestamps in the inflated answers.
Workaround
Download the object and use files.upload instead of register_files.
- Lingua principale
- Python
- Stelle
- 4k
- Fork
- 1k
- Merge medio
- 1g 18h
- PR unite (30g)
- 65
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di googleapis/python-genai
-
Curated history keeps half a user turn when the model turn is invalidForse già presa @Venkaiahbabuneelam l’ha presa 4 giorni fa. Apertapriority: p2 status:awaiting user response type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
googleapis/python-genai#3051 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
[Bug]: Unsubscripted typing.List and typing.Dict crash convert_if_exist_pydantic_model in AFCForse già presa @Venkaiahbabuneelam l’ha presa 4 giorni fa. Apertapriority: p2 type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
googleapis/python-genai#3044 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Name collision in google.genai.interactions: triggers.Interaction shadows response model Interaction in static type checkersForse già presa @Venkaiahbabuneelam l’ha presa 11 giorni fa. Apertapriority: p2 type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
googleapis/python-genai#3013 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
googleapis/python-genai#3078 ·
I maintainer di solito rispondono entro 1 giorno
-
Callable tools: optional scalar parameters (`int | None`) get `"type": "object"` in `parameters_json_schema`, so Gemini sends `{}`Forse già presa @Venkaiahbabuneelam l’ha presa 3 giorni fa. Apertapriority: p2 type: bug
googleapis/python-genai#3060 · 1 commento · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di googleapis/python-genai
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
NVIDIA/earth2studio#1241 ·
I maintainer di solito rispondono entro 3 giorni
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0Apertadocumentation priority: low size: S
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
I maintainer di solito rispondono entro 1 giorno
-
bug help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
I maintainer di solito rispondono entro 1 giorno
-
Broken link in index.rstApertadocumentation
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 65/100
ansys/pydpf-core#3547 ·
I maintainer di solito rispondono entro 1 giorno
-
good first issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
OktoLabsAI/okto-pulse#114 ·
I maintainer di solito rispondono entro 1 giorno