Gemini 3.5 Transcribe Sets New Speech-to-Text Standard
Google has introduced Gemini 3.5 Transcribe, a new speech-to-text model designed to make voice interactions more accurate and natural. It turns spoken language into clean, well-formatted text, removes filler words, handles self-corrections and recognizes specialized terminology. Developers can already test it in preview through Google AI Studio and the Gemini Enterprise Agent Platform.
The new model can be used for voice agents, real-time captioning and post-call analysis. For real-time applications, audio can be streamed through the Live API with sub-second latency, while the Interactions API supports recordings, timestamps and speaker identification. The model automatically detects more than 85 languages and can work with custom vocabularies.
Gemini 3.5 Transcribe also marks a significant improvement in accuracy. According to Artificial Analysis, its WER is 4.0% for streaming and 2.6% for non-streaming use, while final transcriptions are generated 70% faster than with Chirp 3. The new model also performs reliably with noisy audio and accurately captures details such as codes and identifiers.
Gemini 3.5 Transcribe is already used across Google products, including Rambler for Gboard and Gemini for macOS. On Android, it can also edit dictated text by voice, including correcting mistakes and changing the writing style. Google also plans to bring voice typing to Chrome, allowing users to dictate text directly into web pages.
