New web-search leaderboard compares 24 AI models on accuracy, cost and latency
Its creators say they made their internal evaluation method public, testing models with the same search setup, with and without web access.
TLDR
In a September 15 launch announcement, the creators describe a leaderboard meant to help people choose a model for search-heavy work. They say the evaluation tests 24 frontier models end to end, putting cost and latency alongside accuracy. The tests use the same search setup and compare performance with and without web access.
Combined views
703
1 Source, first seen 15d ago