Report
Every seeks someone to lead an AI model benchmarking product
Every’s head of evals says they’re working with 15 people internally to build their own personal benchmarks.
TLDR
Every’s head of evals says CEO Dan Shipper asked for a way to quantify “vibe checks” of new frontier AI models. They’re working with 15 people internally to build personal benchmarks, and the new hire would own the product’s direction. The post describes an ideal candidate as someone who values both quantitative and qualitative judgment.
Combined views
10.7K
5 Sources, first seen ago
