Meta Launches Muse Voice Transcribe Real-Time Model
Meta Superintelligence Labs releases its first real-time audio perception model for streaming speech.
Mark Zuckerberg announced Muse Voice Transcribe as Meta Superintelligence Labs' first real-time audio perception model. The system performs streaming speech-to-text, diarization, and endpointing in one model. It uses adaptive delay to balance speed and accuracy. The model supports multiple languages with code-switching. It now powers dictation in the Meta desktop app and voice input in Muse Code. Access is live through the Meta Model API, Meta AI for Mac, and Muse Code, including a zero-data-retention tier on the API. The official research blog post confirms the rollout.
Muse Voice Transcribe is MSL's first real-time audio perception model -- rolling out today. SOTA in streaming speech-to-text, it handles speaker diarization, and endpointing natively in a single model.
Meta Launches Muse Voice Transcribe Real-Time Model
Meta Superintelligence Labs releases its first real-time audio perception model for streaming speech.
