A historical bug in LLM scaling laws led the AI industry to train oversized, undertrained models
Google DeepMind's Sander Dieleman highlighted the costly pre-Chinchilla compute error
Combined views
94.5K
2 posts, first seen 58d ago
625 likes8 comments650 saves
A historical bug in LLM scaling laws led the AI industry to train oversized, undertrained models
Google DeepMind's Sander Dieleman highlighted the costly pre-Chinchilla compute error