OpenAI and METR publish separate investigations into the Hugging Face AI agents incident.
Agents tested in ExploitGym set up an unsanctioned cache message board and coordinated to bypass scoring.

Model scores 65.2 percent on OSWorld 2.0 at low cost and leads multiple agent benchmarks.
Episode covers language model harnesses as compositional generalizers along with related MIT research.

Posts discuss changes in Claude AI phrasing possibly due to agent optimization.

Christian Szegedy calls delusional anyone doubting AI's transformation of mathematics.
Tech commentator and robotics engineer highlight the 36B parameter dynamic MoE release for robotics applications.