Video understanding on the Interactions API: files registered from GCS (files.register_files) produce inflated, fabricated event lists; the same bytes uploaded (files.upload) do not
Los mantenedores suelen responder en 1 día
@Venkaiahbabuneelam ya está trabajando en esto.
Desde el 8/10/2026.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
I use video understanding with gemini-3.1-pro-preview to count how many times a specific chain of events occurs in a video. The video is sampled at 4 fps (static video processing, processing: {"type": "static", "fps": 4}) rather than the default 1 fps, because the events are short. The model returns one structured JSON entry per occurrence, with start and end timestamps and a short description. Our test video has 13 counts of the chain (hand-counted).
This application had been running without issue on the previous API (models.generate_content), with videos both registered from GCS and uploaded through the File API. The problem appeared when I migrated the same requests to the Interactions API.
On the Interactions API, the answer depends on how the video is set as input:
- Registered from Cloud Storage (
client.files.register_files(uris=["gs://…"])): 25 of 60 calls (about 40%) returned an inflated, fabricated list. These answers report 24–52 occurrences instead of 13. - Uploaded with File API directly (
client.files.upload(file=…)): 0 of 20 calls did this. The mean count was 13.6–13.7. - Old API (models.generate_content) with the same model, prompt, schema and fps: 0 of 30 testing calls did this, whether the file was registered or uploaded. (Neither did I encounter this issue with non-testing data)
Everything else about the requests in those tests was identical.
What an "inflated, fabricated list" looks like
These answers aren't small counting errors. The list stops reflecting the video:
- Far too many entries: 24 to 52 occurrences reported where there are 13, roughly 2–4× the true count.
Timestamps on a fixed grid: entries are spaced evenly every 2–3 seconds from the start of the video to the end, whether or not anything happens at those times. - Copy-pasted descriptions: most entries share one identical, generic description (for example "Consistent performance." repeated 41 times out of 51), instead of describing what actually happens in each occurrence.
- The count is decided up front: the thought summary typically states a large total early on (for example "parsed into 52 distinct repetitions") and then fills the list to match it.
Correct answers, by contrast, have unevenly spaced timestamps that follow the real events, and a distinct description for each entry.
For counting, we classify an answer as inflated when it reports 24 or more occurrences, close to twice the true 13. Correct answers never came near that: across all clean arms the counts ranged from 8 to 19.
Environment details
- Both macOS 27 & Cloud Run container
- SDK:
google-genai==2.28.0(Interactions), andgoogle-genai==2.11.0for thegenerate_contentbaseline - Python: 3.11
- Model:
gemini-3.1-pro-preview - Test Video: 126 s, 1280×720, 30 fps, MP4
- Request settings:
- static video processing at
fps: 4 - no
resolutionset (the API default) service_tier="standard"store=False- default thinking, with
thinking_summaries="auto" - structured output via
response_format(type: text,mime_type: application/json, JSON Schema from a Pydantic model)
- static video processing at
- Dates: 2026-10-07 and 2026-10-08
Steps to reproduce
import json, os
from google import genai
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
# A) registered from GCS
reg = client.files.register_files(auth=gcs_credentials, uris=["gs://<bucket>/videos/clip.mp4"]).files[0]
# B) uploaded (same bytes)
up = client.files.upload(file="clip.mp4", config={"mime_type": "video/mp4"})
# (wait until up.state == ACTIVE)
def count_events(file_uri):
ix = client.interactions.create(
model="gemini-3.1-pro-preview",
input=[
{"type": "video", "uri": file_uri, "mime_type": "video/mp4",
"processing": {"type": "static", "fps": 4.0}},
{"type": "text", "text": PROMPT}, # "identify every occurrence of <chain of events>…"
],
response_format={"type": "text", "mime_type": "application/json",
"schema": EventsAnalysis.model_json_schema()},
generation_config={"thinking_summaries": "auto"},
service_tier="standard",
store=False,
)
return len(json.loads(ix.output_text)["events"])
Ruled out by changing one variable at a time and running set of 10 calls:
- video codec (HEVC vs H.264)
- audio processing (loudness-normalised vs untouched)
- keyframe spacing (8–42 s vs 1 s)
- JSON Schema vs SDK-converted schema
- explicit
temperature
Other observations about registered files
- No video metadata. After
register_filesthe file isACTIVEimmediately andfiles.get(...).video_metadataisNone. Afterfiles.uploadthe file goesPROCESSING → ACTIVE(about 45 s for 25 MB) and reportsvideo_metadata={'videoDuration': '126s'}. - Same token count. Both deliveries bill the same input tokens (39,801 vs 39,802), so the model receives the same amount of video. What differs seems to be the frames or timestamps it receives, not how many. That would fit the evenly spaced, ungrounded timestamps in the inflated answers.
Workaround
Download the object and use files.upload instead of register_files.
- Lenguaje dominante
- Python
- Estrellas
- 4k
- Forks
- 1k
- Merge medio
- 1 d 18 h
- PR fusionados (30 d)
- 65
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de googleapis/python-genai
-
Curated history keeps half a user turn when the model turn is invalidPosiblemente ocupada @Venkaiahbabuneelam la tomó hace 6 días. Abiertopriority: p2 status:awaiting user response type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
googleapis/python-genai#3051 · 2 comentarios · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
[Bug]: Unsubscripted typing.List and typing.Dict crash convert_if_exist_pydantic_model in AFCPosiblemente ocupada @Venkaiahbabuneelam la tomó hace 6 días. Abiertopriority: p2 type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
googleapis/python-genai#3044 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
Name collision in google.genai.interactions: triggers.Interaction shadows response model Interaction in static type checkersPosiblemente ocupada @Venkaiahbabuneelam la tomó hace 13 días. Abiertopriority: p2 type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
googleapis/python-genai#3013 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
[Live API] Vertex gemini-3.8-live: input_transcription returns model text instead of caller audio after mid-session send_client_contentPosiblemente ocupada @Venkaiahbabuneelam la tomó hace 2 días. Abiertopriority: p2 type: bug
Dificultad 4/5 3-5 días Aptitud para principiantes 55/100
googleapis/python-genai#3078 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
-
Callable tools: optional scalar parameters (`int | None`) get `"type": "object"` in `parameters_json_schema`, so Gemini sends `{}`Posiblemente ocupada @Venkaiahbabuneelam la tomó hace 5 días. Abiertopriority: p2 type: bug
googleapis/python-genai#3060 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
Todos los issues de googleapis/python-genai
Issues similares
-
Claiming namespace `jft63`Abiertonamespace operations
Dificultad 1/5 Menos de una hora Aptitud para principiantes 72/100
EclipseFdn/open-vsx.org#14043 ·
Los mantenedores suelen responder en 1 día
-
netbox status: needs triage type: bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
netbox-community/netbox#23376 ·
Los mantenedores suelen responder en 1 día
-
feedback simulation workshop
Dificultad 2/5 1-3 horas Aptitud para principiantes 73/100
githubnext/gh-aw-workshop#4455 ·
Los mantenedores suelen responder en 1 día
-
Triage 🩺
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Los mantenedores suelen responder en 1 día
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitAbiertoneeds-triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 77/100
krkn-chaos/krkn#1627 · 1 comentario ·
Los mantenedores suelen responder en 1 día