Who should grade AI? a16z describes the case for independent testing
Vals AI CEO Rayan Krishnan argues for continuously evolving evaluations as public benchmarks saturate and models get better at optimizing for tests, a16z says.
Vals AI CEO Rayan Krishnan argues for continuously evolving evaluations as public benchmarks saturate and models get better at optimizing for tests, a16z says.
a16z says AI model capability is still mostly self-reported. It describes a discussion in which Vals AI CEO Rayan Krishnan makes the case for independent, continuously evolving evaluations. Other topics include measuring a model’s ability to improve itself, why good benchmarks eventually need retiring, who should set the rules for models, and what happens when token spending begins to rival employee salaries.
173K
2 posts, first seen 3d ago
Vals AI CEO Rayan Krishnan argues for continuously evolving evaluations as public benchmarks saturate and models get better at optimizing for tests, a16z says.
a16z says AI model capability is still mostly self-reported. It describes a discussion in which Vals AI CEO Rayan Krishnan makes the case for independent, continuously evolving evaluations. Other topics include measuring a model’s ability to improve itself, why good benchmarks eventually need retiring, who should set the rules for models, and what happens when token spending begins to rival employee salaries.
Not enough discussion yet.
No sentiment analysis available yet.
Not enough discussion yet.
No sentiment analysis available yet.
—
Not ranked yet
—
Not ranked yet