• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI
Report

Language models trained with simulated faults reportedly grow more resilient as they scale

A post says MIT researchers trained models with simulated faults, finding greater resilience at sizes up to 930 million parameters.

1 Source, 27m ago, first seen 27m ago

TLDR

A post describes an October 7, 2026, MIT preprint testing Llama-2-style models with four-element blocks randomly zeroed during training and inference. It says models trained with these faults grew more resilient as they scaled to roughly 930 million parameters, while models trained without them collapsed when exposed to the same faults. The authors conjecture this approach could enable inference on lower-energy, less reliable chips, but its performance beyond about a billion parameters remains untested.

Combined views

—

1 Source, first seen 27m ago

— likes— comments— saves— reposts

Combined views

—

1 Source, first seen 27m ago

— likes— comments— saves— reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Featured Source

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

Brian Roemmele@BrianRoemmeleReliable Minds From Unreliable Parts: A Forgotten 1956 Von Neumann Paper, And The MIT Discovery That AI Grows Sturdier On Faulty Chips. In 1956 John von Neumann published the lectures he had given at Caltech four years earlier under the title “Probabilistic Logics and the Synthesis of Reliable Organisms from Unreliable Components.” He asked how a machine of any useful size could produce correct outputs when every one of its parts failed with some positive probability. The brain supplied the existence proof. Neurons are noisy, their connections are imprecise, and large regions can be lost without extinguishing coherent thought. The same problem in digital logic, he argued, required deliberate redundancy and restoration at every stage so that error would not accumulate. Once transistors could be made almost perfect, the practical question receded. High voltage and careful process control suppressed the variability that would otherwise flip a gate; the energy cost scaled with the square of that voltage. The architectures that followed assumed the underlying arithmetic would be exact. On October 7, 2026 a preprint from MIT’s Department of Electrical Engineering and Computer Science, the McGovern Institute, and the Department of Physics reopened the problem for language models. Trevor McCourt, Ila R. Fiete, and Isaac L. Chuang trained Llama-2-style transformers on a 350-billion-token slice of FineWeb while randomly zeroing blocks of four matrix elements inside the attention and feed-forward multiplications with probability \(p\). The same stochastic block drops were present at inference. Across roughly 40,000 GPU-hours and models up to \(9.3 \times 10^8\) parameters, the models trained under these faults grew more resilient as they scaled; models trained without them collapsed once the same faults appeared at inference. The authors fit a modified neural scaling law containing an extra curvature term that is absent when \(p = 0\). That term accounts for the recovery of capacity at larger \(N\). They interpret the result as evidence that the networks learn to compute inside “good” error-correcting codes whose relative overhead remains finite even as the model grows without bound. The conjecture is that appropriately trained language models may therefore be formally fault-tolerant. If the conjecture holds, inference could move onto low-voltage digital, analog, or other deliberately imperfect accelerators whose energy per operation is far lower than that of the high-reliability parts used today. The experiments stop near one billion parameters. The authors note that confirming the asymptotic behavior at commercial scales would require something on the order of \(10^8\) GPU-hours, comparable to the training budget of a recent 405-billion-parameter model, and they make the trained checkpoints available for inspection. Almost no public discussion of the result has appeared in the three days since the preprint went up. The practical path the work sketches is therefore still open: train once under the faults the target hardware will produce, then run the resulting model on chips that trade absolute reliability for lower energy. The limit is the untested regime beyond a billion parameters; the action is the experiment that would settle whether the learned correction continues to improve.27m
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    1 Source

    Brian Roemmele@BrianRoemmeleReliable Minds From Unreliable Parts: A Forgotten 1956 Von Neumann Paper, And The MIT Discovery That AI Grows Sturdier On Faulty Chips. In 1956 John von Neumann published the lectures he had given at Caltech four years earlier under the title “Probabilistic Logics and the Synthesis of Reliable Organisms from Unreliable Components.” He asked how a machine of any useful size could produce correct outputs when every one of its parts failed with some positive probability. The brain supplied the existence proof. Neurons are noisy, their connections are imprecise, and large regions can be lost without extinguishing coherent thought. The same problem in digital logic, he argued, required deliberate redundancy and restoration at every stage so that error would not accumulate. Once transistors could be made almost perfect, the practical question receded. High voltage and careful process control suppressed the variability that would otherwise flip a gate; the energy cost scaled with the square of that voltage. The architectures that followed assumed the underlying arithmetic would be exact. On October 7, 2026 a preprint from MIT’s Department of Electrical Engineering and Computer Science, the McGovern Institute, and the Department of Physics reopened the problem for language models. Trevor McCourt, Ila R. Fiete, and Isaac L. Chuang trained Llama-2-style transformers on a 350-billion-token slice of FineWeb while randomly zeroing blocks of four matrix elements inside the attention and feed-forward multiplications with probability \(p\). The same stochastic block drops were present at inference. Across roughly 40,000 GPU-hours and models up to \(9.3 \times 10^8\) parameters, the models trained under these faults grew more resilient as they scaled; models trained without them collapsed once the same faults appeared at inference. The authors fit a modified neural scaling law containing an extra curvature term that is absent when \(p = 0\). That term accounts for the recovery of capacity at larger \(N\). They interpret the result as evidence that the networks learn to compute inside “good” error-correcting codes whose relative overhead remains finite even as the model grows without bound. The conjecture is that appropriately trained language models may therefore be formally fault-tolerant. If the conjecture holds, inference could move onto low-voltage digital, analog, or other deliberately imperfect accelerators whose energy per operation is far lower than that of the high-reliability parts used today. The experiments stop near one billion parameters. The authors note that confirming the asymptotic behavior at commercial scales would require something on the order of \(10^8\) GPU-hours, comparable to the training budget of a recent 405-billion-parameter model, and they make the trained checkpoints available for inspection. Almost no public discussion of the result has appeared in the three days since the preprint went up. The practical path the work sketches is therefore still open: train once under the faults the target hardware will produce, then run the resulting model on chips that trade absolute reliability for lower energy. The limit is the untested regime beyond a billion parameters; the action is the experiment that would settle whether the learned correction continues to improve.27m
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet