/TechMeta's Muse Spark 1.1 scores 48.5 on RadLE 2.0 radiology benchmarkThe model closely approaches the human expert baseline score of 52.0AWJRRNXMRGJWZMCombined views122.6K8 posts, first seen 49d ago1.3K likes commentsMeta's Muse Spark 1.1 scores 48.5 on RadLE 2.0 radiology benchmarkThe model closely approaches the human expert baseline score of 52.0AWJRRNXMRGJWZMReactions7 postsAWAlexandr Wang@alexandr_wang7 weeks agoMuse Spark 1.1 is SOTA on the Radiologists Last Exam Handover Readiness Index (RadLE-H), nearing human expert performance. https://twitter.com/drdatta_aiims/status/2076689436092326276likes: 504replies: 48bookmarks: 45reposts: 40JRJack Rae@jack_w_rae7 weeks agoMuse Spark 1.1 is really strong on health topics! I have found this to be the case from personal usage but it’s cool to see it show up on benchmarks also. https://twitter.com/alexandr_wang/status/2076696459005837410likes: 75replies: 8bookmarks: 6reposts: 2ZMZvi Mowshowitz@TheZvi7 weeks agoLove it but also laughing at the idea of a Last Exam 2.0. https://twitter.com/DrDatta_AIIMS/status/2076689408040849507likes: 408replies: 20bookmarks: 27reposts: 15RNRamez Naam@ramez7 weeks ago@TheZvi Last exam v 132.7likes: 2replies: 0bookmarks: 0reposts: 0RGRiley Goodside@goodside7 weeks ago@TheZvi Reverse nominative determinism; same reason there’s lots of Final Fantasy sequels.likes: 11replies: 0bookmarks: 0reposts: 0JWJason Wei@_jasonwei7 weeks agoMuse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans are still a lot better, but we are working on closing the gap! https://twitter.com/DrDatta_AIIMS/status/2076689408040849507likes: 267replies: 27bookmarks: 32reposts: 32Show allCombined views122.6K8 posts, first seen 49d ago1.3K likes
AWAlexandr Wang@alexandr_wang7 weeks agoMuse Spark 1.1 is SOTA on the Radiologists Last Exam Handover Readiness Index (RadLE-H), nearing human expert performance. https://twitter.com/drdatta_aiims/status/2076689436092326276likes: 504replies: 48bookmarks: 45reposts: 40
JRJack Rae@jack_w_rae7 weeks agoMuse Spark 1.1 is really strong on health topics! I have found this to be the case from personal usage but it’s cool to see it show up on benchmarks also. https://twitter.com/alexandr_wang/status/2076696459005837410likes: 75replies: 8bookmarks: 6reposts: 2
ZMZvi Mowshowitz@TheZvi7 weeks agoLove it but also laughing at the idea of a Last Exam 2.0. https://twitter.com/DrDatta_AIIMS/status/2076689408040849507likes: 408replies: 20bookmarks: 27reposts: 15
RGRiley Goodside@goodside7 weeks ago@TheZvi Reverse nominative determinism; same reason there’s lots of Final Fantasy sequels.likes: 11replies: 0bookmarks: 0reposts: 0
JWJason Wei@_jasonwei7 weeks agoMuse Spark 1.1 outperforms GPT-5.6 Sol and Gemini 3.1 on Radiology's Last Exam. We don't beat Fable (yet). And Humans are still a lot better, but we are working on closing the gap! https://twitter.com/DrDatta_AIIMS/status/2076689408040849507likes: 267replies: 27bookmarks: 32reposts: 32