Google launches Gemini 3.8 Live, beating OpenAI and xAI in key voice benchmarks

Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two new real-time voice models built for conversations that can combine speech, camera input, reasoning and external tools.

The standard Gemini 3.8 Live is designed for large-scale deployment and lower-cost use. It can process live camera input, switch between 97 supported languages even mid-sentence, and call external tools through the API in the background without interrupting the conversation. Extended Thinking is aimed at more complex tasks and can continue speaking while it reasons or waits for tools, reducing the long pauses common in voice assistants.

In third-party testing by Artificial Analysis, Gemini 3.8 Live Extended Thinking narrowly outperformed comparable models from OpenAI and xAI across several voice and agentic benchmarks:

  • Speech-to-Speech Quality Index: 82.6, compared with 81.5 for GPT-Live-1 Astra and 81.3 for Grok Voice Think Fast 2.0.
  • τ-Voice: 68.6%, versus 67.9% for GPT-Live-1 Astra and 56.5% for Grok Voice.
  • τ³-Banking: 35.1%, compared with 32.0% for GPT-Live-1 Astra and 16.5% for xAI-Realtime.
  • Big Bench Audio: 97.7% accuracy for Gemini 3.8 Live Extended Thinking.

Google is also competing aggressively on price. Artificial Analysis estimates one hour of input audio at about $3.50 for Extended Thinking and $0.84 for Gemini 3.8 Live, compared with $4.80 for Grok Voice Think Fast 2.0 and $5.83 for GPT-Live-1 Astra.

Both models are available through Google AI Studio and the Gemini Live API. Google has also begun bringing them to its own products, with Gemini 3.8 Live appearing in Search Live and Extended Thinking rolling out across Gemini Live and selected Workspace services.