• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Muse Voice Transcribe Launched by MSL

    First streaming audio model from Meta Superintelligence Lab does real-time ASR with diarization.

    AA
    AW
    LB
    31 Sources, 29d ago, first seen 29d ago

    TLDR

    Bowen Cheng announced the launch of Muse Voice Transcribe from Meta Superintelligence Lab. The model processes audio in 80ms chunks and decides at each step whether to keep listening or output text. It supports hour-long sessions, over 20 speakers, multilingual input with code-switching, and contextual biasing. Meta AI accounts and Mark Zuckerberg posted that it reaches the speed-accuracy pareto frontier on streaming transcription. The model is rolling out now to developers and appears in the Meta AI macOS app.

    Combined views

    323.8K

    31 Sources, first seen 29d ago

    Combined views

    323.8K

    31 Sources, first seen 29d ago

    2.2K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2.2K likes
    130 comments
    362 saves
    1.1K reposts
    130 comments
    362 saves
    1.1K reposts

    31 Sources

    @finkdIt holds up on messy, real audio too -- trained across 70+ languages (with 25 validated at launch), handles mid-sentence code-switching, and manages hour-long sessions with 20+ speakers. You can bias it toward names and domain terms as well.
    @AIatMetaWith adaptive delay, Muse Voice Transcribe achieves the pareto frontier on speed-accuracy trade-off measured by time to final transcription.
    @bowenc0221Muse can hear now - we launched Muse Voice Transcribe today! It is the first streaming audio perception model from MSL which does ASR, diarization and endpointing all in real-time. It supports hour-long audio, 20+ speakers, multilingual, seamless code-switching, and contextual biasing. We shared detailed design and what makes Muse Voice Transcribe different from other real-time voice models in our blog post. We give full control on when and how long to listen to audio back to the model. And with RL, it learns “adaptive delay” to wait longer only on a few hard words. This is one of the key contributions to achieving the new pareto frontier on speed/accuracy trade-off. Muse Voice Transcribe is available today via Meta Model API, Meta AI for Mac, and Muse Code. Please give it a try and let us know how you like it: https://research.meta.ai/blog/introducing-muse-voice-transcribe
    @alexandr_wang2/ it’s accurate and quick - it uses adaptive delay to decide how much context it needs before transcribing each word. trained on 70+ languages (with 25 validated at launch), handles code-switching mid-sentence, and manages hour+ sessions and 20+ speakers without post-processing.
    @imhaotianRT @bowenc0221: Muse can hear now - we launched Muse Voice Transcribe today! It is the first streaming audio perception model from MSL whic…
    @rohanpaul_aiLove this, another huge release from Meta. Lunched Muse Voice Transcribe for real-time voice dictation, with the lowest 3.1% final-transcription word error rate with adaptive delay. which is significantly ahead of competing models. - The big deal is that Muse does something beyond just transcribing speech; it learns when to wait, when to commit a word, when a speaker changes, and when a turn is actually over, all inside the same streaming model. That makes it much closer to a real-time perception layer for voice agents than a conventional speech-to-text API. - Muse processes audio in 80ms chunks and chooses after each chunk whether to emit text or keep listening. Meta made it available through Meta Model API, Meta AI for Mac, and Muse Code
    @spencerbarnettIntroducing Muse Voice Transcribe! 🗣️ This is our first real-time audio perception model from @AIatMeta. It's trained on 70+ languages and delivers diarization with 20+ speakers. Available now in the Meta AI macOS app!
    @delipraowhatever your perspectives on Meta’s AI capabilities are, you have to concede there’s a new pipeline in place that’s churning out interesting work at an impressive clip.
    @ArtificialAnlysOn First Partial Transcript, Muse Voice Transcribe achieves a 3.6% WER at 0.13s after end of speech, just ahead of ElevenLabs Scribe v2 Realtime on accuracy and latency. Muse is more accurate than Cartesia Ink-2 with external endpoints at 4.0% WER and 0.07s, while trading some speed, and is both more accurate and faster than AssemblyAI Universal-3.5 Pro Realtime, whose Max Accuracy and Min Latency modes achieve 4.0% WER at 0.18s and 0.17s, respectively.
    @JackKMuse Voice Transcribe: another model from MSL available to developers & agent builders that pushes the pareto frontier... this time coming in at the top mark. Try it at https://dev.meta.ai

    Sentiment

    Positive89.2%10.8%Negative

    Summary

    Sentiment

    Positive89.2%10.8%Negative

    Many accounts praised Muse Voice Transcribe for its speed, accuracy, diarization, and Franglais support, while some called it useless or obsolete compared to multimodal models.

    Based on 92 sentiment-bearing replies from 74 accounts across 5 conversations.

    Summary

    Many accounts praised Muse Voice Transcribe for its speed, accuracy, diarization, and Franglais support, while some called it useless or obsolete compared to multimodal models.

    Based on 92 sentiment-bearing replies from 74 accounts across 5 conversations.

    31 Sources

    @finkdIt holds up on messy, real audio too -- trained across 70+ languages (with 25 validated at launch), handles mid-sentence code-switching, and manages hour-long sessions with 20+ speakers. You can bias it toward names and domain terms as well.
    @AIatMetaWith adaptive delay, Muse Voice Transcribe achieves the pareto frontier on speed-accuracy trade-off measured by time to final transcription.
    @bowenc0221Muse can hear now - we launched Muse Voice Transcribe today! It is the first streaming audio perception model from MSL which does ASR, diarization and endpointing all in real-time. It supports hour-long audio, 20+ speakers, multilingual, seamless code-switching, and contextual biasing. We shared detailed design and what makes Muse Voice Transcribe different from other real-time voice models in our blog post. We give full control on when and how long to listen to audio back to the model. And with RL, it learns “adaptive delay” to wait longer only on a few hard words. This is one of the key contributions to achieving the new pareto frontier on speed/accuracy trade-off. Muse Voice Transcribe is available today via Meta Model API, Meta AI for Mac, and Muse Code. Please give it a try and let us know how you like it: https://research.meta.ai/blog/introducing-muse-voice-transcribe
    @alexandr_wang2/ it’s accurate and quick - it uses adaptive delay to decide how much context it needs before transcribing each word. trained on 70+ languages (with 25 validated at launch), handles code-switching mid-sentence, and manages hour+ sessions and 20+ speakers without post-processing.
    @imhaotianRT @bowenc0221: Muse can hear now - we launched Muse Voice Transcribe today! It is the first streaming audio perception model from MSL whic…
    @rohanpaul_aiLove this, another huge release from Meta. Lunched Muse Voice Transcribe for real-time voice dictation, with the lowest 3.1% final-transcription word error rate with adaptive delay. which is significantly ahead of competing models. - The big deal is that Muse does something beyond just transcribing speech; it learns when to wait, when to commit a word, when a speaker changes, and when a turn is actually over, all inside the same streaming model. That makes it much closer to a real-time perception layer for voice agents than a conventional speech-to-text API. - Muse processes audio in 80ms chunks and chooses after each chunk whether to emit text or keep listening. Meta made it available through Meta Model API, Meta AI for Mac, and Muse Code
    @spencerbarnettIntroducing Muse Voice Transcribe! 🗣️ This is our first real-time audio perception model from @AIatMeta. It's trained on 70+ languages and delivers diarization with 20+ speakers. Available now in the Meta AI macOS app!
    @delipraowhatever your perspectives on Meta’s AI capabilities are, you have to concede there’s a new pipeline in place that’s churning out interesting work at an impressive clip.
    @ArtificialAnlysOn First Partial Transcript, Muse Voice Transcribe achieves a 3.6% WER at 0.13s after end of speech, just ahead of ElevenLabs Scribe v2 Realtime on accuracy and latency. Muse is more accurate than Cartesia Ink-2 with external endpoints at 4.0% WER and 0.07s, while trading some speed, and is both more accurate and faster than AssemblyAI Universal-3.5 Pro Realtime, whose Max Accuracy and Min Latency modes achieve 4.0% WER at 0.18s and 0.17s, respectively.
    @JackKMuse Voice Transcribe: another model from MSL available to developers & agent builders that pushes the pareto frontier... this time coming in at the top mark. Try it at https://dev.meta.ai