MIT Professor Challenges Pure DNNs on Scaling Behavior
MIT professor Omar Khattab states frontier models add non-differentiable parts to standard networks.
Omar Khattab posted that current frontier systems are not pure differentiable DNNs. He noted they include scratchpads, offloaded context, PTC, and recursion, which let them exceed plain GPT-4 with RLHF on functional tasks. Susan Zhang shared a screenshot of Yann LeCun quoting an earlier critique of LLMs as an AGI offramp. Other replies discussed whether past critics have shifted positions without acknowledgment. Khattab urged keeping hypotheses clear between conventional DNNs and augmented setups. The thread stays on technical definitions rather than broader claims.
Nicely, this is entirely falsifable. You think DNNs with differentiable operations are all you need? Fantastic - just show the same scaling and generalization behavior with one. Even on purely functional tasks that don’t inherently involve a tool env. Only text in, text out.