Best frontier AI agent reportedly passes barely more than a third of database benchmark queries
Weaviate Podcast describes a benchmark built by Shreya Shankar and colleagues with 54 queries spanning 12 datasets and four database systems, testing how AI agents answer questions across databases.
TLDR
Weaviate Podcast says Shreya Shankar counted 25-plus data-agent benchmarks and found none that tested querying multiple databases. It describes the Data Agent Benchmark she built with colleagues at UC Berkeley, Washington University and Hasura to address that gap: 54 queries spanning 12 datasets and four database systems. According to the podcast, the best frontier AI agent passes barely more than a third of the queries. It also describes agents avoiding unfamiliar SQL dialects by exporting tables to flat files and doing the work in Pandas instead.
Combined views
10.6K
4 Sources, first seen 16d ago