A deeper bench for AI evaluation beyond METR
The post says strict nondisclosure agreements keep many evaluators from discussing their work publicly, even as they contribute to AI system and model documentation.
TLDR
A user calls for more talent at existing AI evaluators, new organizations across areas of expertise, and stronger standards and frameworks for frontier AI auditing. The post argues that voluntary commitments are not enough: standards, regulation, and frameworks could give evaluators more consistent, durable operating conditions.
Combined views
447.4K
49 Sources, first seen 23d ago