Google DeepMind released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, two text-to-speech models featuring voice replication and support for 100 languages. The Flash TTS model scored 89.5% on the Pronunciation Robustness Benchmark, taking the top spot from the previous leader, Gemini 3.1 Flash TTS. It also debuted at #2 on the Provider Voice Arena Leaderboard with an Elo of 1,263, while the Flash-Lite version ranked #6 with 1,236. The models are available via the Gemini API and AI Studio with a library of over 2,000 voices. Gemini 3.8 Flash TTS costs $32.98 per 1 million characters, and Flash-Lite is $22.07 per 1 million characters, both priced higher than the $18.31 rate for Gemini 3.1 Flash TTS but below Eleven v3 at $100 per 1 million characters. In terms of speed, the Flash model processes 44.1 characters per second, roughly 2.7 times faster than real-time. However, it trails other competitors such as Falcon 2, which processes 204.9 characters per second, and Luna TTS at 153.5 characters per second.

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
  2. SOURCEhuggingnewshuggingnews.com
  3. SOURCEhuggingnewshuggingnews.com
  4. SOURCEmarketbrief.now
Markdown