Ex-OpenAI VP Jerry Tworek Declares Era of AI Evals Over
Reactions from ranked influencers
2 postsFrontier Notes from AGI House: Jerry Tworek @MillionInt on why “the era of evals is done.” 𝐅𝐮𝐥𝐥 𝐢𝐧𝐭𝐞𝐫𝐯𝐢𝐞𝐰 ↓ Jerry Tworek spent seven years at OpenAI, where he worked on research that helped shape modern coding and reasoning models, including Codex and reinforcement learning for reasoning. He later served as VP of Research, and is now CEO of @CoreAutoAI, building toward automated AI research. At AGI House, Jerry sat down with @rockyrmit for a technical conversation on Codex, HumanEval, Copilot, RL, tool use, test-time compute, and automated AI labs. We cover: › Why code became the first serious LLM vertical › What HumanEval taught the field about verifiable rewards › Why Jerry thinks “the era of evals is done” › How real-world deployment differs from static benchmarks › What GitHub Copilot taught OpenAI about product quality › How reinforcement learning shaped modern reasoning behaviors › Why tool use is “99% systems and 1% algorithms” › Whether test-time compute still has room to scale › What it takes to build automated AI research labs Watch the full interview below.
Combined views
3.6K
2 posts, first seen 3h ago