Ai2 Builds BenchMIRT to Audit LLM Evaluation Benchmarks
Ai2 tool audits whether benchmark questions test intended model abilities such as safety.
The Allen Institute for AI announced BenchMIRT to audit LLM safety and capability evaluations and determine which model abilities their questions actually test. On the BBQ social-bias benchmark the tool showed questions distinguished models more by reasoning ability than by safety. The institute posted the details in a thread that includes a diagram and links to their blog post on the project.
Combined views
14.5K
3 posts, first seen 1d ago
