Meta Shares Muse Spark 1.2 Multimodal Benchmarks
Meta AI Research posted benchmarks and demos for the model's visual and robotics tasks.
Meta AI Research published a blog post on the multimodal intelligence of Muse Spark 1.2. Company posts describe performance across visual reasoning and chart understanding benchmarks plus code generation from image or video inputs. A variant serves as a robot brain that decodes instructions, calls tools, observes outcomes, and iterates until tasks finish. Demos illustrate bimanual robot planning and agentic media generation. Alexandr Wang and AI at Meta accounts posted the updates along with photos and videos.
1/ muse spark 1.2 is a very strong multimodal model—it can do visual coding, robotics planning, and audio-visual understanding that all come together through agentic tools.