---
format: "aidr-story-markdown/v1"
id: "56559347616b53b6365a8fa772244c9bb51bdd75a033a26dcd799d59627622b0"
canonical_url: "https://aidr.today/56559347?lang=en"
title: "Meta Launches First Real Time Audio Perception Model Muse Voice Transcribe for $3 per 1,000 Minutes"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-01T18:51:47.000Z"
category: "Models"
topics: ["meta","llm","inference","benchmark"]
source_urls: ["https://huggingnews.com/ai/meta-launches-first-real-time-audio-perception-model-muse-voice-transcri-c183cb9b","https://x.com/finkd/status/2094836602681938385","https://x.com/AIatMeta/status/2094839236016976028","https://x.com/AIatMeta/status/2094839238495801457","https://x.com/finkd/status/2094836605517234402","https://x.com/finkd/status/2094836606834262023","https://x.com/AIatMeta/status/2094839240748138978"]
summary: "Meta Superintelligence Labs released Muse Voice Transcribe, a multimodal LLM that combines streaming automatic speech recognition and speaker diarization. The model processes audio in 80ms chunks and employs adaptive delay to commit words based on complexity, targeting a 3.1% final transcription word error rate. It supports more than 70 languages, including 25 validated at launch, and can manage hour long sessions with over 20 speakers. The tool is available for $3 per 1,000 audio minutes through the Meta model API, which includes a zero data retention tier. Muse Voice Transcribe currently powers voice input in Muse Code and dictation in the Meta desktop app. It ranks first on Artificial Analysis streaming speech to text benchmarks and public diarization tests."
---

# Meta Launches First Real Time Audio Perception Model Muse Voice Transcribe for $3 per 1,000 Minutes

> [Open the canonical story](<https://aidr.today/56559347?lang=en>)

**Published:** 2026-09-01T18:51:47.000Z
**Category:** Models
**Topics:** meta, llm, inference, benchmark

## Summary

Meta Superintelligence Labs released Muse Voice Transcribe, a multimodal LLM that combines streaming automatic speech recognition and speaker diarization\. The model processes audio in 80ms chunks and employs adaptive delay to commit words based on complexity, targeting a 3\.1% final transcription word error rate\. It supports more than 70 languages, including 25 validated at launch, and can manage hour long sessions with over 20 speakers\. The tool is available for $3 per 1,000 audio minutes through the Meta model API, which includes a zero data retention tier\. Muse Voice Transcribe currently powers voice input in Muse Code and dictation in the Meta desktop app\. It ranks first on Artificial Analysis streaming speech to text benchmarks and public diarization tests\.

## Sources

- [Story source](<https://huggingnews.com/ai/meta-launches-first-real-time-audio-perception-model-muse-voice-transcri-c183cb9b>)
- [Story source](<https://x.com/finkd/status/2094836602681938385>)
- [Story source](<https://x.com/AIatMeta/status/2094839236016976028>)
- [Supporting source](<https://x.com/AIatMeta/status/2094839238495801457>)
- [Supporting source](<https://x.com/finkd/status/2094836605517234402>)
- [Story source](<https://x.com/finkd/status/2094836606834262023>)
- [Supporting source](<https://x.com/AIatMeta/status/2094839240748138978>)

