Who should grade AI? a16z describes the case for independent testing
Vals AI CEO Rayan Krishnan argues for continuously evolving evaluations as public benchmarks saturate and models get better at optimizing for tests, a16z says.
TLDR
a16z says AI model capability is still mostly self-reported. It describes a discussion in which Vals AI CEO Rayan Krishnan makes the case for independent, continuously evolving evaluations. Other topics include measuring a model’s ability to improve itself, why good benchmarks eventually need retiring, who should set the rules for models, and what happens when token spending begins to rival employee salaries.
Combined views
205.7K
3 Sources, first seen ago