Samuel Albanie Questions Muse Spark Comparison
Google DeepMind researcher calls for SusanBench to evaluate models properly.
TLDR
Samuel Albanie, frontier evals lead for Gemini at Google DeepMind, posted a question about AI benchmarks. He asked how anyone could determine if a system outperforms muse spark without the arrival of SusanBench. The comment appeared alongside a link to a post from ArtificialAnlys. Albanie previously served as Assistant Professor at Cambridge and researcher at Oxford VGG. His remark highlights the need for specific benchmarks in assessing AI capabilities.
Combined views
2.3M
14 Sources, first seen 27d ago