By Jess Weatherbed
Publication Date: 2026-08-26 17:00:00
Google says that 3.5 Transcribe “represents a major advancement from our previous transcription model, Chirp 3,” especially regarding multilingual performance and wording error rates. The transcription model allows users to “edit naturally with just your voice,” according to Google, and can automatically format text and remove filler words like “um” and “uh.”
Users can provide a customized vocabulary to the model, allowing 3.5 Transcribe to automatically adapt transcription to unique spelling requirements and specialized jargon to prevent those words from being edited manually. It can also attribute speech for up to three speakers in pre-recorded audio, alongside providing word-level timestamps.
Alongside 3.5 Transcribe, Google also said that 3.5 Live and 3.5 Live Experimental updates will be coming to Gemini Audio today that build on the existing speech recognition tech powering Gemini’s voice chat mode. After we published this story, Google then reached out to say…


