Google DeepMind Pilots Double-Blind AI Evaluations
The company creates secure environments so evaluators cannot identify the models under review.
Google DeepMind stated it is piloting double-blind evaluations for frontier AI. The method creates a secure environment where neither team knows the identities of the models being assessed. Nat Purser highlighted what he described as the first such evaluation of a proprietary model, involving Google DeepMind and Averi. Demis Hassabis reposted the announcement from the company account. The goal mentioned is building trust in benchmarks for proprietary models through cryptographic security.
Combined views
8.2K
2 posts, first seen 1d ago
Google DeepMind Pilots Double-Blind AI Evaluations
The company creates secure environments so evaluators cannot identify the models under review.
Google DeepMind stated it is piloting double-blind evaluations for frontier AI. The method creates a secure environment where neither team knows the identities of the models being assessed. Nat Purser highlighted what he described as the first such evaluation of a proprietary model, involving Google DeepMind and Averi. Demis Hassabis reposted the announcement from the company account. The goal mentioned is building trust in benchmarks for proprietary models through cryptographic security.