Google rolls out Gemini 3.8 Live and 3.8 Live Extended Thinking with parallel reasoning for voice AI

Google has introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new live dialogue models. The announcement covers their capabilities, performance results, developer support, audio watermarking, and availability.

Gemini 3.8 Live and Live Extended Thinking

Google says the two models bring upgrades in intelligence and parallel reasoning for voice interactions and complex tasks. Gemini 3.8 Live is built for scale and cost efficiency, combining conversational intelligence, fluid dialogue, and visual grounding.

Meanwhile, Gemini 3.8 Live Extended Thinking is built for high-complexity tasks, with increased intelligence and multi-step reasoning. For developers and enterprises, the models provide building blocks for reliable, production-ready voice agents.

The models also support voice interactions with Gemini across the Gemini app, Google Workspace, and Search, allowing users to tackle complex tasks using voice.

Performance

Gemini 3.8 Live Extended Thinking recorded the following results in the benchmarks cited by Google:

  • 82.6 on Artificial Analysis Speech to Speech Quality Index, securing the #1 overall spot
  • 68.6% on τ-Voice for agentic task completion
  • 35.1% on Sierra’s τ-Voice-banking benchmark for agentic task completion
  • 97.7% on Big Bench Audio

Google also says Gemini 3.8 Live Extended Thinking maintains a competitive price point compared with other frontier models. Meanwhile, Gemini 3.8 Live secured second place in the Speech Agent Arena.

On ServiceNow’s EVA-Bench, a benchmark for evaluating voice agents, Google says the models push the Pareto Frontier for complex workflows by balancing accuracy with conversational quality.

Note: This was run on the Live API on Gemini Enterprise Agent Platform.
Voice and visual capabilities

Gemini 3.8 Live processes visual inputs in near real time, adding visual context to conversations. It can also automatically detect and transition between 97 supported languages during a conversation.

The model can execute tools and API calls in the background while continuing the conversation. It can acknowledge requests and keep chatting while those tasks are completed.

For more complex workflows, Gemini 3.8 Live Extended Thinking can reason and speak simultaneously. It uses early verbal cues such as “Let me check that…” to acknowledge prompts and provides live progress narration while completing multi-step background tasks.

Across Google Workspace and Search, the Live models also support voice interactions for complex tasks.

Developer and enterprise ecosystem

The Gemini Live API supports developer platforms including:

  • Agora
  • Fishjam
  • LangChain
  • LiveKit
  • Pipecat
  • Vercel
  • Vision Agents

These platforms manage complex real-time media streaming infrastructure behind the scenes, allowing developers to focus on the user experience.

Google is also working with Salesforce, Genspark, and Lumeris, which have highlighted the models latency, fluidity, and tool-calling capabilities.

SynthID watermarking

All audio generated by Google AI products is watermarked with SynthID. The imperceptible watermark is woven directly into the audio output, allowing AI-generated content to remain detectable and helping prevent misinformation.

Google has also published a model card covering safety and responsibility details.

Availability

Gemini 3.8 Live is available through:

  • Developers: Gemini API and Google AI Studio
  • Enterprises: Private preview in Gemini Enterprise; coming soon to Gemini Enterprise for Customer Experience
  • Everyone: Search Live

Gemini 3.8 Live Extended Thinking is available through:

  • Developers: Gemini API and Google AI Studio
  • Enterprises: Private preview in Gemini Enterprise; coming soon to Gemini Enterprise for Customer Experience and Google Workspace business customers
  • Everyone: Gemini Live; Google AI Pro and Ultra subscribers in Workspace in Docs; all Google AI subscribers in Gmail and Keep


Related Post