Google Gemini 3.5 Transcribe: 85 Languages, Real-Time Slip Correction, 4% WER
Google has unveiled Gemini 3.5 Transcribe, a speech-to-text model that recognizes over 85 languages and automatically removes filler words like 'um' and 'uh' while correcting slips of the tongue in real time. It achieves a 4.0 percent word error rate in streaming mode, with 70 percent lower latency than its predecessor, Chirp 3. Through function calling, the model can hand off tasks to other Gemini models, enabling more complex workflows. This marks a significant leap in transcription accuracy and speed, positioning Google to challenge established players like OpenAI's Whisper and AssemblyAI.