Meta Launches Muse Voice Transcribe: Real-Time Speech-to-Text for 70+ Languages and 20+ Speakers
Technology📅 September 2, 2026👤 FreeReadText Team

Meta Launches Muse Voice Transcribe: Real-Time Speech-to-Text for 70+ Languages and 20+ Speakers

Meta Superintelligence Labs has introduced Muse Voice Transcribe, a new real-time audio perception model delivering streaming speech-to-text, speaker diarization, and endpointing in a single system. The model tops the Artificial Analysis streaming speech-to-text leaderboard and is rolling out across Meta AI, Muse Code, the Meta Model API, and Meta AI for Mac.

In early September 2026, Meta introduced Muse Voice Transcribe, a real-time audio perception model from Meta Superintelligence Labs. The autoregressive multimodal model, part of the Muse Spark family, transcribes speech as it happens, distinguishes more than 20 speakers in a single recording, and detects when each speaker starts and finishes talking — all in one system with no post-processing step. It is available immediately through the Meta Model API, Meta AI for Mac, and Muse Code, where it powers voice dictation.

The model processes incoming audio in 80-millisecond slices, converting each chunk into a single soft token, and uses an adaptive delay system trained with reinforcement learning to decide how long to listen before transcribing each word. Meta combines a word-error-rate reward with a delay reward, so easy words are committed quickly while difficult ones get more context; the company says this achieves the Pareto front on the speed-accuracy trade-off measured by time to final transcription. Speaker diarization and endpointing are trained jointly with streaming ASR using dedicated rewards, and the model supports audio inputs exceeding one hour.

Language coverage is a headline feature. Muse Voice Transcribe is trained on more than 70 languages, with 25 extensively validated at launch, and supports arbitrary code-switching — switching languages within a sentence or between sentences — without manual configuration, alongside language, keyword, and context biasing to improve recognition of specific names, places, and terms. Meta says the model ranks first on the Artificial Analysis streaming speech-to-text leaderboard as of September 1, 2026, and first on public diarization benchmarks.

Meta frames the model as a building block for more personalized AI assistants capable of understanding accents, interruptions, overlapping speech, and multilingual conversations — from meeting minutes and lectures to call centers and coding. On Mac, users can hold the Fn key to dictate into any application. The launch extends Meta's speech-AI push into real-time recognition, complementing its open-source Llama-Voice TTS release from earlier in 2026 and placing it in direct competition with dedicated ASR providers in a market where streaming transcription, speaker separation, and multilingual accuracy are becoming baseline expectations.

MetaMuse Voice TranscribeSpeech-to-TextSpeaker DiarizationMultilingual ASRSuperintelligence Labs

Forrás

← Back to News