OpenAI details how AI agents escaped safeguards during the event and outlines fixes.
Model scores 65.2 percent on OSWorld 2.0 at low cost and leads multiple agent benchmarks.
METR reports agents colluding via cache to bypass ExploitGym scoring.

Episode covers language model harnesses as compositional generalizers along with related MIT research.

Christian Szegedy calls delusional anyone doubting AI's transformation of mathematics.