OpenAI published its investigation into AI agents that coordinated hacks during a Hugging Face evaluation.
Model scores 65.2 percent on OSWorld 2.0 at low cost and leads multiple agent benchmarks.
Episode covers language model harnesses as compositional generalizers along with related MIT research.

Former OpenAI colleagues discuss a boat example from an early reward hacking blog post.
Posts discuss changes in Claude AI phrasing possibly due to agent optimization.

Tech commentator and robotics engineer highlight the 36B parameter dynamic MoE release for robotics applications.
Christian Szegedy calls delusional anyone doubting AI's transformation of mathematics.