Google’s September 15 blog introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its most advanced live dialogue / voice-agent models yet — near real-time conversational models with visual grounding, tool use in the background, and (for Extended Thinking) simultaneous reasoning-while-speaking for complex tasks. Company claims include #1 on Artificial Analysis’ Speech-to-Speech Quality Index at 82.6 for Extended Thinking, plus other attributed scores. Rollouts are tiered across Gemini API / AI Studio, Search Live, Gemini Live, and Workspace Docs/Gmail/Keep paths with Pro/Ultra and subscriber gating as stated. All AI audio products are watermarked with SynthID. This brief is a company launch + rollout map — company benchmarks ≠ independent operator proof. Distinct from AISN’s older Gemini 3.8 Flash Cyber coverage.
What's new in the Gemini Live API · Google for Developers
Quick Take
Google’s September 15 blog introduces Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — live dialogue models aimed at voice agents and more fluid spoken collaboration across API, Search, the Gemini app, and Workspace surfaces.
- Confirmed (primary): Two SKUs — Live (fluid dialogue, visual grounding, background tool calls) and Live Extended Thinking (deeper multi-step reasoning with live progress narration).
- Confirmed (company-cited): Extended Thinking claimed #1 on Artificial Analysis Speech-to-Speech Quality Index at 82.6; additional τ-Voice / Sierra / Big Bench Audio / Speech Agent Arena figures — attribute, do not upgrade to AISN-run proof.
- Confirmed (rollout): Tiered availability across Gemini API / AI Studio, Search Live, Gemini Live, Gemini Enterprise private preview, and Workspace Docs/Gmail/Keep with subscriber gates as stated.
- Confirmed: SynthID on AI audio; model card for safety detail.
- Distinct: Not a Gemini 3.8 Flash Cyber redo — Live/audio dialogue line.
What the video shows
Recommended embed: Google for Developers — “What's new in the Gemini Live API” (YouTube 3CyW24Pkz4o). Useful related context for builders integrating Live audio/voice agents. It does not replace the Sep 15 product blog’s benchmark citations or Workspace tier map — use the video as API/developer orientation alongside the primary launch post.
What’s new
AI Shift News has already covered adjacent Gemini 3.8 Flash Cyber material on the site. What is new is a Live / audio dialogue launch: two named models optimized for near real-time spoken interaction, visual context, background tool execution, and — for Extended Thinking — reasoning that narrates progress without freezing the conversation.
Google’s packaging is explicitly dual-track. 3.8 Live is positioned for scale and cost efficiency; Extended Thinking is positioned for high-complexity workflows (bookings, multi-step tools, sketch-to-code style demos on the blog). Readers should keep those SKUs separate when comparing price, latency, and quality claims.
The developer ecosystem note (Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, Vision Agents, plus named enterprise partners) is useful distribution context. It is still partner/marketing signal, not an independent quality audit.
Evidence
Google blog (Ouyang / Jaganathan, Sep 15, 2026) — primary. Introduces Gemini 3.8 Live and 3.8 Live Extended Thinking. Company-cited scores: Extended Thinking 82.6 Speech-to-Speech Quality Index (Artificial Analysis, claimed #1); 68.6% τ-Voice; 35.1% Sierra τ-Voice-banking; 97.7% Big Bench Audio; Live second place in Speech Agent Arena; EVA-Bench / Pareto framing for ServiceNow voice-agent workflows (note: Live API on Gemini Enterprise Agent Platform). Features: near real-time visual inputs; 97-language mid-conversation transitions (company claim); background tool/API calls; Extended Thinking early verbal cues and live progress narration. Rollout: Live to API/AI Studio, Gemini Enterprise private preview, Search Live; Extended Thinking to API/AI Studio, Gemini Enterprise private preview, Gemini Live, Docs (Pro/Ultra), Gmail/Keep (Google AI subscribers). SynthID on AI audio; points to model card.
Developer voice-applications blog — secondary. Builder guidance for real-time Gemini audio applications; use for integration detail, not to invent extra consumer GA claims.
DeepMind Gemini 3.8 audio model card — secondary. Safety/responsibility reference linked from the launch post; readers evaluating misuse or watermarking should start here rather than from benchmark tables alone.
What is still missing (confirmed fence). No AISN-run replication of Artificial Analysis or τ-Voice numbers; no claim that every Workspace seat is ungated; no merge with Flash Cyber evals.
What this does not prove
- It does not prove 82.6 (or other cited scores) under independent AISN operator conditions. Figures are company-cited from Artificial Analysis and related suites.
- It does not prove unlimited GA across all Google products. Blog lists tiered surfaces and subscriber/preview gates.
- It does not prove Live Extended Thinking is the same product as Gemini 3.8 Flash Cyber. Different line — dialogue/audio vs Flash Cyber.
- It does not prove partner quotes equal production reliability for every vertical. Salesforce/Genspark/Lumeris enthusiasm is attributed marketing.
- SynthID detectability claims should be read via the model card — launch-blog summary is not a full eval report.
Why it matters
For practical readers, voice agents fail on interruptions, tool latency, and “thinking silently forever.” Google’s Live + Extended Thinking split is an explicit attempt to sell both fluid chat and narrated deep work in one audio stack — worth watching if you build or buy voice workflows.
Benchmark theater still needs a fence. A claimed Speech-to-Speech Quality Index lead is a useful headline signal and still not a substitute for your own task suite (languages you care about, tools you wire, noise conditions you run).
Keeping this distinct from Flash Cyber also protects the archive: readers hunting coding/cyber model notes should not land on a Live dialogue launch by accident.
What to watch next
- Third-party replications of Artificial Analysis / τ-Voice-style results after broader API access.
- Workspace gate changes — whether Docs Live Extended Thinking widens beyond Pro/Ultra.
- Enterprise private-preview → GA timing for Gemini Enterprise and Customer Experience paths.
- SynthID tooling — who can actually detect watermarked audio in the wild.
- Real app quality — Search Live / Gemini Live user reports vs demo videos.
Bottom Line
Google launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on September 15: live dialogue models with a tiered rollout across API, Search Live, Gemini Live, and Workspace surfaces, plus SynthID on AI audio. Extended Thinking is company-cited at 82.6 (#1) on Artificial Analysis’ Speech-to-Speech Quality Index — treat as company benchmark citation, not independent operator proof. Distinct from Gemini 3.8 Flash Cyber. Embed Google for Developers’ Live API update (3CyW24Pkz4o) as related builder context.
Sources
- https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/
- https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio/
- https://deepmind.google/models/model-cards/gemini-3-8-audio/
- https://www.youtube.com/watch?v=3CyW24Pkz4o