Meta Launches First Real Time Audio Perception Model Muse Voice Transcribe for $3 per 1,000 Minutes
Models30d agoMeta Superintelligence Labs released Muse Voice Transcribe, a multimodal LLM that combines streaming automatic speech recognition and speaker diarization. The model processes audio in 80ms chunks and employs adaptive delay to commit words based on complexity, targeting a 3.1% final transcription word error rate. It supports more than 70 languages, including 25 validated at launch, and can manage hour long sessions with over 20 speakers.
The tool is available for $3 per 1,000 audio minutes through the Meta model API, which includes a zero data retention tier. Muse Voice Transcribe currently powers voice input in Muse Code and dictation in the Meta desktop app. It ranks first on Artificial Analysis streaming speech to text benchmarks and public diarization tests.
Key sources
- SOURCE@finkd“SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model”x.com
- SOURCE@aiatmeta“ranks first on @ArtificialAnlys streaming speech-to-text and on public diarization benchmarks”x.com
- SUPPORT@aiatmeta“Audio is processed in 80ms chunks (12.5 Hz), one token each”x.com
- SUPPORT@finkd“trained across 70+ languages (with 25 validated at launch)”x.com
- SOURCE@finkd“Live now on the Meta model API with a zero-data-retention tier”x.com
- SUPPORT@aiatmeta“achieves the pareto frontier on speed-accuracy trade-off measured by time to final transcription”x.com