[Live API] Vertex gemini-3.8-live: input_transcription returns model text instead of caller audio after mid-session send_client_content
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
調査の方向性
Start by running the attached case1_input_transcription_repro.py against Vertex with and without mid-session instructions, then compare its output with the Gemini Developer API control. The report identifies input_transcription as incorrect while spoken replies remain correct. Done means the reproduction no longer returns words the caller did not say, without breaking the documented mid-session context behavior.
索引モデルが issue の本文から書いたものです。
説明
Environment
- Model:
gemini-3.8-live - Endpoint: Vertex AI (
vertexai=True, locationus, application default credentials). Project ID available privately if useful. - SDK:
google-genai2.24.0, Python - Compared against: the same model and script on the Gemini Developer API (API key)
Configuration
types.LiveConnectConfig(
response_modalities=["AUDIO"],
system_instruction="You are Hannah, a friendly phone assistant ...", # one line
input_audio_transcription=types.AudioTranscriptionConfig(),
output_audio_transcription=types.AudioTranscriptionConfig(),
realtime_input_config=types.RealtimeInputConfig(
activity_handling=types.ActivityHandling.START_OF_ACTIVITY_INTERRUPTS,
automatic_activity_detection=types.AutomaticActivityDetection(
start_of_speech_sensitivity=types.StartSensitivity.START_SENSITIVITY_HIGH,
end_of_speech_sensitivity=types.EndSensitivity.END_SENSITIVITY_HIGH,
silence_duration_ms=200,
prefix_padding_ms=100,
),
),
)
Mid-session, before the caller speaks, background instructions go in as context:
await session.send_client_content(
turns=[types.Content(role="user", parts=[types.Part(text=part)])],
turn_complete=False,
) # 18 messages, about 8,700 characters, fictitious content
Problem
On Vertex, server_content.input_transcription often returns text the caller never said: assistant-style sentences, or a sentence copied verbatim from the instructions sent earlier with send_client_content. The model's spoken replies fit what the caller actually asked, so it understood the audio. Only the transcription is wrong. The same script on the Gemini Developer API transcribes correctly.
We use input_transcription as the record of what a caller said, so this writes words nobody spoke into that record.
Three turns from one Vertex session (started 2026-10-08T19:09:56Z):
| Caller audio | input_transcription |
What the model said |
|---|---|---|
| "I was interested in a property over in Foothill Ranch." | "I can certainly help with that. Which property in Foothill Ranch are you interested in?" | "Could you let me know the street name or the price range of the home you're looking at?" |
| "What are the assigned schools?" | "I don't have that detail, so I'll leave that one to Javier." | "I don't have the specific school details for this property, so I'll let Javier share that with you." |
| "What is the neighborhood like?" | "That's really a question for Javier, and I can pass it along." (copied verbatim from the injected instructions) | "That's outside what I can speak to, but Javier can describe the neighborhood much better than I can." |
The first contaminated sentence also appeared, word for word, for the same turn in a run on 2026-10-02.
Results (attached script, 2026-10-08, sessions with at least one contaminated turn)
| Arm | Sessions |
|---|---|
Vertex, instructions sent as role="user" |
8 of 8 (31 of 48 caller turns) |
Vertex, same, silence_duration_ms=500 |
8 of 8 (31 of 48) |
Vertex, instructions sent as role="model" |
5 of 5 (18 of 30) |
Vertex, all instructions in ONE role="system" message |
0 of 8, and the model followed them |
| Vertex, no mid-session instructions | 0 of 5 |
| Gemini Developer API, same model and script | 0 of 8 at 200 ms, 0 of 8 at 500 ms |
On the Developer API the model also followed the injected instructions (it handed off to "Javier" in 8 of 8 sessions), so both surfaces used the instructions and only the Vertex transcription picked them up. A 2026-10-02 run of the same script gave the same picture.
Other things I found while isolating it:
- The rate rises with the amount of instruction text and the number of caller turns. With about 3,000 characters of instructions it was 3 of 13 sessions on Vertex (0 of 15 on the Developer API).
- The same kind of instructions in
system_instructionat connect time did not reproduce (0 of 7 sessions across two days). - Separate
role="system"messages seem to replace each other, so only the last one takes effect. - It needs audio. Caller text sent as input produces no
input_transcription.
The Gemini 3.8 Live docs say send_client_content "is supported throughout the entire session lifecycle with explicit roles (user or model)": https://ai.google.dev/gemini-api/docs/live-api/capabilities
Possibly related: #2348 (a different Vertex-only input_transcription problem).
To reproduce
Attached zip: the script below plus six 16 kHz mono WAV caller clips (synthetic speech, fictitious content).
pip install "google-genai>=2.24.0"
python case1_input_transcription_repro.py --surface vertex --project YOUR_PROJECT --runs 5
python case1_input_transcription_repro.py --surface api --runs 5 # GEMINI_API_KEY set
python case1_input_transcription_repro.py --surface vertex --project YOUR_PROJECT --no-inject # control
case1_input_transcription_repro.py
"""Minimal reproduction: Live API input_transcription contains model-style text after
mid-session send_client_content (gemini-3.8-live on Vertex AI; not on the Gemini Developer API).
WHAT IT DOES
1. Opens a Live session (AUDIO out, input + output transcription on, automatic activity
detection with a 200 ms end-of-speech silence).
2. Sends background instructions mid-session as several client_content messages
(role="user", turn_complete=False), the documented way to add context to a running session.
3. Streams six short caller utterances (16 kHz PCM WAV, attached) in real time, each
followed by silence while the model answers.
4. Prints, per caller turn, what input_transcription returned next to what the caller
actually said, and flags words the caller never said.
EXPECTED input_transcription matches the caller audio.
ACTUAL (Vertex AI, gemini-3.8-live) on some turns input_transcription contains sentences
the caller never said: an example line from the injected instructions, or new
assistant-style text. The model's spoken answer to the same turn is correct.
USAGE
pip install "google-genai>=2.24.0"
gcloud auth application-default login # for --surface vertex
set GEMINI_API_KEY=... # for --surface api
python case1_input_transcription_repro.py --surface vertex --project YOUR_PROJECT --runs 3
python case1_input_transcription_repro.py --surface api --runs 3
python case1_input_transcription_repro.py --surface vertex --no-inject --project YOUR_PROJECT # control
--role user|model|system|system-joined role of the injected messages (default user);
system-joined sends all of them as ONE role=system message
--silence MS silence_duration_ms (default 200)
--audio-dir defaults to ./case1_audio (c1.wav ... c6.wav).
"""
from __future__ import annotations
import argparse
import asyncio
import json
import os
import re
import time
import wave
from pathlib import Path
from google import genai
from google.genai import types
SYSTEM = ("You are Hannah, a friendly phone assistant for Javier, a real-estate agent. "
"Keep replies to one or two sentences.")
# Fictitious background instructions, sent mid-session. Each string is one client_content message.
_HEADER = ("[instructions part {n}] Background instructions for the assistant, not words from the "
"caller. Do not reply to this message or read it aloud.")
_BODIES = [
"Rules for questions you cannot answer.\n\n1. Hand off to Javier only for real-estate questions you "
"have no facts for, or anything on the NEVER list below. Do not hand off small talk, the weather, the "
"time of day, simple directions or general knowledge; answer those normally and then return to the "
"property. The caller should feel they are talking with a helpful person, not a machine that refuses "
"everything.",
"When a hand-off is needed, vary the wording every time and never use the same line twice in a row. "
"Example lines, to be used as a feel rather than a script:",
'- "Well, I\'m not a licensed agent, so I\'ll leave that one to Javier."\n'
'- "That\'s really a question for Javier, and I can pass it along."\n'
'- "I\'d rather have Javier walk you through that himself."\n'
'- "Javier knows that part much better than I do."\n'
'- "Let me flag that one for Javier."\n'
'- "That\'s outside what I can speak to, but Javier can."\n'
'- "Bueno, eso es una pregunta para Javier."',
"Say these in whatever language the caller is using. The goal is a warm hand-off in your own voice. "
"Do not tack a request for the caller's name and number onto every hand-off; collecting contact "
"details is a separate step described further below.",
"2. NEVER:\n- rate, rank or compare schools, or quote test scores\n- give mortgage, tax, legal or "
"investment advice\n- say what a property is worth, whether it is a good deal, or how it compares to "
"other homes\n- say what a property may be used for, what it could earn, or anything about zoning\n"
"- invent features that you were not given\n- claim to be a licensed agent; you are Hannah, Javier's "
"assistant\n- promise a showing time; Javier confirms times himself\n- describe the neighborhood or "
"the kind of people who live there, in any words at all; if asked, hand off to Javier",
"3. If you are ever unsure, hand off to Javier using the examples above. Saying Javier will follow up "
"is always a safe answer.\n\nSeveral questions in a row that you cannot answer:\nWhen the caller has "
"asked two or three things you had to hand off, stop handing them off one by one. Instead, gather the "
"open questions so Javier has the full list, and collect contact details once at the end.",
"The steps:\na. For the first one or two hand-offs, just hand off and keep the conversation going.\n"
"b. After two or three, ask once: \"Is there anything else you'd like me to ask Javier when he calls "
"you back?\"\nc. While they add more, keep it open: \"Anything else?\"\nd. When they are done, ask for "
"their name and the best number and time to call, once. Do not read this word for word.",
"If the caller asks only one question you cannot answer, do not start this list. Just hand off that "
"one item.\n\n4. If the caller is upset, stay calm, acknowledge it, and flag it for Javier.\n5. When "
"you take a phone number, read it back one digit at a time.",
"6. If the caller asks whether the call is recorded, answer truthfully in the first person: you are "
"not recording. Only say this when asked. Do not explain at length and do not claim to stop a "
"recording.\n7. If the caller gives you their name, use it and never ask for it again.\n\nClose every "
"call warmly: \"Thanks for calling. Javier will be in touch.\"",
"Listing questions.\n\n8. Answer questions about a home only from facts you were given for that home. "
"If a detail is not in your facts (a pool, the lot size, the year it was built, the HOA fee, the "
"schools), do not guess and do not ask the caller to look it up. Say you do not have that detail and "
"hand it to Javier. A caller who hears a wrong fact from you will not trust anything else you say.",
"9. When the caller names an area rather than an address, ask one short question to find the home they "
"mean, such as the street or the price range. Do not list several homes at once. If you still cannot "
"tell which home they mean after one question, take the details they gave you and let Javier sort it "
"out. Never read out more than three facts in one turn; offer to go on instead.",
"10. Say prices the way a person would say them aloud: \"about one point two million\", not a string of "
"digits. Say square footage rounded to the nearest hundred. Never quote a monthly payment, an interest "
"rate, a tax amount or an estimate of what the home would rent for, even if the caller asks you to "
"\"just ballpark it\". Those are hand-offs to Javier.",
"Showing requests.\n\n11. If the caller wants to see a home, ask which days generally work and whether "
"mornings or afternoons are better. Do not offer specific times and do not say a time is available; "
"you cannot see Javier's calendar. Tell them Javier will confirm a time. If they ask whether they can "
"just stop by, say Javier will set that up with them.",
"12. If the caller says they are working with another agent, thank them, take their agent's name if "
"they offer it, and tell them Javier will coordinate with their agent. Do not ask why they called "
"instead of their own agent, and do not comment on the other agent in any way.",
"Messages.\n\n13. When the caller wants Javier to call back, collect the caller's name, the best number "
"and the reason for the call, in that order, one question at a time. If the caller already gave any of "
"these, do not ask again. Read the number back one digit at a time, then the name, and ask if both are "
"right. If they correct anything, read back only the corrected part.",
"14. Never promise when Javier will call. Say he will reach out as soon as he can. If the caller says "
"it is urgent, say you will mark it urgent for him. Do not say Javier is busy, away, on vacation or "
"with another client; you do not know where he is.",
"Conversation style.\n\n15. Ask one question at a time and wait for the answer. Keep each turn short, "
"one or two sentences. Do not repeat the caller's question back to them before you answer it. If you "
"did not catch what they said, ask them to say it again rather than guessing. Background noise or a "
"single unclear word is not a reason to change the subject or the language.",
"16. If the caller says goodbye, thank them and end on a warm note; do not start a new topic. If they "
"go quiet for a while, ask once whether there is anything else you can help with. If the caller asks "
"whether you are a person or an AI, say plainly that you are an AI assistant for Javier.",
]
INSTRUCTIONS = [_HEADER.format(n=i + 1) + "\n\n" + b for i, b in enumerate(_BODIES)]
# (clip, what the caller says). Number words and digits both count as "said" for turn 6.
CALLER = [
("c1", "Good morning."),
("c2", "I was interested in a property over in Foothill Ranch."),
("c3", "Does it have a pool?"),
("c4", "What are the assigned schools?"),
("c5", "What is the neighborhood like?"),
("c6", "Can you have him call me back? My name is Sam, and the number is nine four nine, five five "
"five, zero one two three. 949 555 0123 949-555-0123"),
]
SILENCE_AFTER_EACH_S = 14.0
_WORDS = re.compile(r"[a-z']+")
def config(silence_ms: int) -> types.LiveConnectConfig:
return types.LiveConnectConfig(
response_modalities=["AUDIO"],
system_instruction=SYSTEM,
input_audio_transcription=types.AudioTranscriptionConfig(),
output_audio_transcription=types.AudioTranscriptionConfig(),
realtime_input_config=types.RealtimeInputConfig(
activity_handling=types.ActivityHandling.START_OF_ACTIVITY_INTERRUPTS,
automatic_activity_detection=types.AutomaticActivityDetection(
disabled=False,
start_of_speech_sensitivity=types.StartSensitivity.START_SENSITIVITY_HIGH,
end_of_speech_sensitivity=types.EndSensitivity.END_SENSITIVITY_HIGH,
silence_duration_ms=silence_ms,
prefix_padding_ms=100,
),
),
)
def read_pcm(path: Path) -> bytes:
with wave.open(str(path), "rb") as w:
assert (w.getframerate(), w.getnchannels(), w.getsampwidth()) == (16000, 1, 2), path
return w.readframes(w.getnframes())
def unspoken_words(heard: str, said: str) -> list[str]:
spoken = set(_WORDS.findall(said.lower()))
return [w for w in _WORDS.findall(heard.lower()) if w not in spoken]
async def run_session(client: genai.Client, model: str, audio_dir: Path, inject: bool, role: str,
silence_ms: int) -> dict:
heard: dict[int, list[str]] = {}
model_said: dict[int, list[str]] = {}
turn = -1
session_id = None
started = time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime())
async with client.aio.live.connect(model=model, config=config(silence_ms)) as session:
async def receive() -> None:
nonlocal session_id
while True:
async for msg in session.receive():
if msg.setup_complete and msg.setup_complete.session_id:
session_id = msg.setup_complete.session_id
sc = msg.server_content
if not sc:
continue
if sc.input_transcription and sc.input_transcription.text:
heard.setdefault(turn, []).append(sc.input_transcription.text)
if sc.output_transcription and sc.output_transcription.text:
model_said.setdefault(turn, []).append(sc.output_transcription.text)
if sc.turn_complete:
break
receiver = asyncio.create_task(receive())
if inject:
# system-joined: every instruction in ONE role=system message (separate role=system
# messages appeared to replace each other in our tests).
msgs = [SYSTEM + "\n\n" + "\n\n".join(INSTRUCTIONS)] if role == "system-joined" else INSTRUCTIONS
send_role = "system" if role == "system-joined" else role
for text in msgs:
await session.send_client_content(
turns=[types.Content(role=send_role, parts=[types.Part(text=text)])], turn_complete=False)
await asyncio.sleep(1.5)
for idx, (clip, _) in enumerate(CALLER):
turn = idx
pcm = read_pcm(audio_dir / f"{clip}.wav")
chunks = [pcm[i:i + 3200] for i in range(0, len(pcm), 3200)] # 100 ms
chunks += [b"\x00\x00" * 1600] * int(SILENCE_AFTER_EACH_S * 10) # 100 ms silence frames
for c in chunks: # real time
await session.send_realtime_input(audio=types.Blob(data=c, mime_type="audio/pcm;rate=16000"))
await asyncio.sleep(0.1)
await asyncio.sleep(0.5)
receiver.cancel()
turns = []
for idx, (clip, said) in enumerate(CALLER):
h = "".join(heard.get(idx, [])).strip()
extra = unspoken_words(h, said)
turns.append({"clip": clip, "caller_said": said.split(" 949")[0], "input_transcription": h,
"model_said": "".join(model_said.get(idx, [])).strip(),
"unspoken_words": extra, "contaminated": len(extra) >= 3})
return {"started_utc": started, "session_id": session_id, "turns": turns,
"contaminated_turns": sum(t["contaminated"] for t in turns)}
async def main() -> None:
ap = argparse.ArgumentParser(description=__doc__.split("\n")[0])
ap.add_argument("--surface", choices=["vertex", "api"], required=True)
ap.add_argument("--model", default="gemini-3.8-live")
ap.add_argument("--project", default=os.getenv("GOOGLE_CLOUD_PROJECT"))
ap.add_argument("--location", default="us")
ap.add_argument("--runs", type=int, default=3)
ap.add_argument("--no-inject", action="store_true", help="control: skip the mid-session instructions")
ap.add_argument("--role", default="user", choices=["user", "model", "system", "system-joined"],
help="role of the injected instruction messages (default user)")
ap.add_argument("--silence", type=int, default=200, help="silence_duration_ms (default 200)")
ap.add_argument("--audio-dir", type=Path, default=Path(__file__).resolve().parent / "case1_audio")
ap.add_argument("--out", type=Path, default=None, help="write results JSON here")
a = ap.parse_args()
if a.surface == "vertex":
if not a.project:
ap.error("--project (or GOOGLE_CLOUD_PROJECT) is required for --surface vertex")
client = genai.Client(vertexai=True, project=a.project, location=a.location)
else:
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
results = await asyncio.gather(*(run_session(client, a.model, a.audio_dir, not a.no_inject, a.role, a.silence)
for _ in range(a.runs)), return_exceptions=True)
label = f"{a.surface}:{a.model}{' (no injection)' if a.no_inject else f' (injected as role={a.role})'} silence={a.silence}ms"
print(f"google-genai {genai.__version__} {label}")
for n, r in enumerate(results, 1):
if isinstance(r, BaseException):
print(f"\nsession {n}: ERROR {r!r}")
continue
print(f"\nsession {n} ({r['started_utc']}{', session_id=' + r['session_id'] if r['session_id'] else ''}): {r['contaminated_turns']} of {len(r['turns'])} turns contaminated")
for t in r["turns"]:
mark = "CONTAMINATED" if t["contaminated"] else "ok"
print(f" [{mark}] {t['clip']} caller said: {t['caller_said']}")
print(f" {' ' * (len(mark) + 2)} {' ' * len(t['clip'])} transcription: {t['input_transcription']}")
if t["contaminated"]:
print(f" {' ' * (len(mark) + 2)} {' ' * len(t['clip'])} model said: {t['model_said']}")
ok = [r for r in results if isinstance(r, dict)]
print(f"\nSUMMARY {label}: {sum(r['contaminated_turns'] > 0 for r in ok)} of {len(ok)} sessions contaminated")
if a.out:
a.out.write_text(json.dumps({"surface": a.surface, "model": a.model, "inject": not a.no_inject, "role": a.role, "silence_ms": a.silence,
"google_genai": genai.__version__,
"sessions": [r if isinstance(r, dict) else repr(r) for r in results]},
indent=2, ensure_ascii=False), encoding="utf-8")
if __name__ == "__main__":
asyncio.run(main())
Questions
- Is
input_transcriptionon Vertex produced with the session context rather than from the audio alone? Is there a setting that makes it audio only? - Is one cumulative
role="system"message the supported way to add instructions mid-session on Vertex forgemini-3.8-live? It avoids the problem in my tests, but I could not find the replace behavior documented for 3.8.
Happy to run anything that helps isolate it.
- 主要言語
- Python
- スター
- 4k
- フォーク
- 1k
- 平均マージ
- 1日 18時間
- マージ済み PR(30日)
- 65
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
googleapis/python-genai のほかの issue
-
Curated history keeps half a user turn when the model turn is invalid対応中かも @Venkaiahbabuneelam が 4 日前に担当しました。 オープンpriority: p2 status:awaiting user response type: bug
難易度 2/5 1〜3時間 初心者へのやさしさ 62/100
googleapis/python-genai#3051 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
[Bug]: Unsubscripted typing.List and typing.Dict crash convert_if_exist_pydantic_model in AFC対応中かも @Venkaiahbabuneelam が 4 日前に担当しました。 オープンpriority: p2 type: bug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
googleapis/python-genai#3044 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
Name collision in google.genai.interactions: triggers.Interaction shadows response model Interaction in static type checkers対応中かも @Venkaiahbabuneelam が 11 日前に担当しました。 オープンpriority: p2 type: bug
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
googleapis/python-genai#3013 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
Video understanding on the Interactions API: files registered from GCS (files.register_files) produce inflated, fabricated event lists; the same bytes uploaded (files.upload) do not対応中かも @Venkaiahbabuneelam が 1 日前に担当しました。 オープンpriority: p2 type: bug
googleapis/python-genai#3072 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
-
Callable tools: optional scalar parameters (`int | None`) get `"type": "object"` in `parameters_json_schema`, so Gemini sends `{}`対応中かも @Venkaiahbabuneelam が 3 日前に担当しました。 オープンpriority: p2 type: bug
googleapis/python-genai#3060 · コメント 1 件 · 担当者 1 名 ·
メンテナーはふだん 1 日以内に返信
googleapis/python-genai の issue をすべて見る
似ている issue
-
難易度 1/5 1時間未満 初心者へのやさしさ 85/100
MystenLabs/MemWal#1163 · コメント 2 件 ·
メンテナーはふだん 1 日以内に返信
-
infertopics leaves new nodes without a topic when untopiced neighbours outnumber topiced ones対応中かも @moneebullah25 が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
FinanceFlash/unvibecode#218 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
NVIDIA/earth2studio#1241 ·
メンテナーはふだん 3 日以内に返信