Language models trained with simulated faults reportedly grow more resilient as they scale
A post says MIT researchers trained models with simulated faults, finding greater resilience at sizes up to 930 million parameters.
TLDR
A post describes an October 7, 2026, MIT preprint testing Llama-2-style models with four-element blocks randomly zeroed during training and inference. It says models trained with these faults grew more resilient as they scaled to roughly 930 million parameters, while models trained without them collapsed when exposed to the same faults. The authors conjecture this approach could enable inference on lower-energy, less reliable chips, but its performance beyond about a billion parameters remains untested.
Combined views
—
1 Source, first seen ago
