Muse Voice Transcribe Launched by MSL
Researcher Bowen Cheng announced the streaming audio model from Meta's Superintelligence Lab.
Bowen Cheng posted the launch of Muse Voice Transcribe from Meta Superintelligence Lab. The autoregressive multimodal model from the Muse Spark family processes streaming audio for automatic speech recognition, diarization, and endpointing. It handles multilingual input with code switching and contextual biasing. Audio arrives in chunks and the model decides at each step whether to keep listening or emit text. Meta's official account described the reinforcement learning approach that produces adaptive delay. Jack Krawczyk stated the model is now available to developers and agent builders through the Meta platform.
Combined views
103.8K
8 posts, first seen 16h ago
