Gemini evaluated with an MLCommons benchmark in an OpenMined secure enclave
A post quoting AVERI’s standards director says Google DeepMind did not see the benchmark, and its owner did not see the Gemini model. The setup allowed the evaluation to produce results.
TLDR
AVERI’s standards director described an evaluation that put a Gemini model and an MLCommons benchmark into an OpenMined secure enclave, with firewalls to protect each side’s material. The director said it built on a UK AISI pilot with Anthropic two years earlier that used a toy model and benchmark.
Combined views
7.5K
2 Sources, first seen 10h ago
