Google Releases Gemini 3.5 Transcribe with 2.6% Average Word Error Rate Across 85+ Languages

Loading…

Google AI has launched Gemini 3.5 Transcribe, a dedicated speech-to-text model reporting a 2.6% average word error rate across more than 85 languages, positioning it as a strong contender in the automatic speech recognition space. The model targets enterprise and developer use cases that require high-accuracy multilingual transcription, including meeting summarization, accessibility tooling, and voice-driven agentic workflows. A 2.6% WER across such a broad language set is a technically significant result, particularly for lower-resource languages where most commercial ASR systems degrade sharply. Developers building voice interfaces or transcription pipelines should benchmark Gemini 3.5 Transcribe against existing solutions like Whisper and competing cloud ASR APIs to assess whether the accuracy gains justify migration. The release expands Google's Gemini product family beyond text and multimodal reasoning into specialized speech understanding.