Shavit Speculates on OpenAI Self-Improvement Pace Versus Anthropic
OpenAI policy lead questions whether reinforcement learning investments create a widening gap.
Yo Shavit posted wondering whether OpenAI’s deeper investment in reinforcement learning could produce faster AI self-improvement than Anthropic, with any gap possibly growing over the next six months. Miles Brundage replied that interpretability investments or work at other labs could change the direction instead. In his own follow-up reply, Shavit added that modeling the situation as an AI race stops making sense immediately before recursive self-improvement because of the world-transforming consequences and high odds of nationalization. Visible replies also covered shifting perceptions of which lab held the lead and differing views on training approaches.
I wonder whether we are starting to see faster AI self-improvement at OpenAI vs. Anthropic based on the former’s known deeper investment in RL and TTC proving more useful for tasks related to AI R&D, and that this gap will grow significantly over the next 6 months. (Obviously…