The 'Bitter Lesson' and the assumptions built into AI training
One post argues that language model training builds in assumptions about data, while reinforcement learning rewards still come mostly from human ingenuity.
TLDR
The post questions how casually the 'Bitter Lesson' gets invoked in AI discussions. It argues that current large language model training introduces inductive bias—built-in assumptions—including that the data distribution is good enough to simply clone. Reinforcement learning faces its own version of the problem, the author says: where do rewards come from? Today, the post argues, the answer is mostly human ingenuity.
Combined views
13.9K
1 Source, first seen 14h ago
The 'Bitter Lesson' and the assumptions built into AI training
One post argues that language model training builds in assumptions about data, while reinforcement learning rewards still come mostly from human ingenuity.
TLDR
The post questions how casually the 'Bitter Lesson' gets invoked in AI discussions. It argues that current large language model training introduces inductive bias—built-in assumptions—including that the data distribution is good enough to simply clone. Reinforcement learning faces its own version of the problem, the author says: where do rewards come from? Today, the post argues, the answer is mostly human ingenuity.