Researchers Announce DiG-Bench for AI Discovery
Benchmark uses interactive text games to test language models on novel concepts.
TLDR
James Whittington announced DiG-bench, a benchmark of text-based discovery games for testing frontier language models. The games operate in the natural domain of language models rather than requiring visual inputs. Ruairidh Battleday described the benchmark as containing new concepts and mechanisms that models must discover by interacting within each game. Whittington stated that the team had tested frontier models and observed their scores. Separate posts from academics including Melanie Mitchell and Tim Rocktäschel shared the announcement and expressed interest in the approach.
Combined views
1.2M
12 Sources, first seen 49d ago
