• btc = $79 500.00 1 148.78 (1.23 %)

  • btc = $79 500.00 1 148.78 (1.23 %)

27 Aug, 2026
1 min time to read

Google is calling it its most accurate speech recognition model yet. It supports more than 85 languages, removes filler words and can distinguish between up to three speakers. Developers can start using it now.

Google has launched Gemini 3.5 Transcribe, a dedicated speech-to-text model that the company describes as its most accurate speech recognition system to date. The model is already being used in the Gemini app for macOS and in Rambler, Gboard’s voice typing feature on Android, in select countries.

In multilingual speech recognition benchmarks shared by Google, Gemini 3.5 Transcribe recorded lower error rates than competing models — with lower scores indicating better performance.

The main difference is that Gemini 3.5 Transcribe goes beyond word-for-word transcription. It can recognize when someone corrects themselves mid-sentence, remove filler words and automatically format the final text. A custom vocabulary helps it handle jargon and unusual spellings, while support extends to more than 85 languages.

For prerecorded audio, the model can also identify up to three different speakers. Google is currently testing support for conversations with more participants.

Gemini 3.5 Transcribe is now available to developers in public preview through the Gemini API in Google AI Studioand Google Antigravity, with enterprise access through the Gemini Enterprise Agent Platform. Google also plans to bring the model to Chrome “soon,” allowing users to dictate into virtually any text field on the web, from a search box to an online form.