Google released its most precise speech-to-text model, Gemini 3.5 Transcribe, which automatically removes filler words and supports more than 85 languages. The tool is integrated into the Gemini app for macOS and the Rambler feature on Android, with developer access provided via Google AI Studio and the Gemini API.

The model achieves a 2.6% word error rate on non-streaming audio and 4% on streaming, reflecting a 70% reduction in final transcription time compared to the previous Chirp 3. It provides speaker attribution for up to three people and offers sub-second bidirectional streaming through the Gemini Live API to handle alphanumeric tokens such as postal codes and order IDs.

Sign in to suggest edits

Key sources

  1. SOURCE@google“removing filler words like "ums" and "ahs" while handling self-corrections”x.com
  2. SOURCE@sundarpichai“API available now in @GoogleAIStudio and Gemini Enterprise, or try it in the Gemini app on macOS or Rambler on Android!”x.com
  3. SUPPORT@officiallogank“new speech to text model with smart transcription, function calling, more precise transcription (lower WER)”x.com
  4. SUPPORT@_philschmid“70% reduction time for the final transcription compared to Chirp 3”x.com
  5. SUPPORT@google“Developers can start building in @GoogleAIStudio and Google @Antigravity”x.com
  6. SUPPORT@google“help turn spoken thoughts into polished text”x.com
Markdown