• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Sakana AI says PC-ALM trains 1,000-layer neural nets without backpropagation

    The lab’s local-learning approach draws on distributed optimization and NeuroAI. It targets a limitation Sakana AI describes in predictive coding: learning signals struggle to reach internal layers in deep networks.

    YL
    DP
    KA
    16 Sources, ,

    TLDR

    Sakana AI has introduced PC-ALM, which it says trains 1,000-layer neural networks using only local dynamics, without backpropagation—the standard approach in deep learning. The method builds on predictive coding, where neuron activations adjust through an energy-minimization process based on local prediction errors. To address predictive coding’s difficulty scaling to deeper networks, the lab says it replaces that energy-based formulation with an augmented Lagrangian, a mathematical tool from distributed optimization.

    Combined views

    381.1K

    16 Sources, first seen 16d ago

    Combined views

    381.1K

    16 Sources, first seen 16d ago

    3.6K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    16d ago
    first seen 16d ago
    3.6K likes
    98 comments
    2.7K saves
    440 reposts
    98 comments
    2.7K saves
    440 reposts

    Sentiment

    Positive73.2%26.8%Negative

    Summary

    Sentiment

    Positive73.2%26.8%Negative

    Many accounts welcomed PC-ALM as an insightful local-learning alternative to backpropagation with potential for spiking nets, while others dismissed the MNIST results as inadequate or unoriginal.

    Based on 44 sentiment-bearing replies from 41 accounts across 4 conversations.

    Summary

    Many accounts welcomed PC-ALM as an insightful local-learning alternative to backpropagation with potential for spiking nets, while others dismissed the MNIST results as inadequate or unoriginal.

    Based on 44 sentiment-bearing replies from 41 accounts across 4 conversations.

    16 Sources

    @SakanaAILabsIntroducing PC-ALM, a local-learning alternative to backpropagation. Our method trains 1000-layer neural nets using only local dynamics, and without backprop. Blog: https://pub.sakana.ai/pc-alm/ Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly. How can a physical system, such as the brain, solve multilayer credit assignment without explicit use of backprop? We look for inspiration in two related fields: distributed optimization and NeuroAI. In NeuroAI, predictive coding asks each neuron activation to solve an energy-based inference problem instead of using a standard forward pass. That inference step can be implemented as energy-minimization dynamics on local prediction errors. This perspective -- each layer as a dynamical system -- has proven promising, but performance of predictive coding hasn't scaled well with depth. Credit signals at far ends of the network struggle to diffuse into internal layers. We turn to distributed optimization, generalizing predictive coding to use an augmented Lagrangian instead of energy. This motivation stems back to a classic 1988 paper by LeCun, showing that the Lagrange multipliers of a deep network can be identified with gradients of a supervised loss. The augmented Lagrangian then bridges LeCun's perspective to the standard predictive coding that is used in NeuroAI. We find that this new perspective yields a natural PC-like alternative to backpropagation, resulting in a method we call PC-ALM. PC-ALM differs from PC in that it introduces dual neurons (Lagrange multipliers) as part of the layer-local dynamics, resulting in each layer acting as a PI feedback control system to minimize local prediction errors. We find that PC-ALM is capable of propagating signals to seemingly arbitrary depth, especially in deep narrow networks where standard PC struggles to learn. Ultimately, our motivation here is to understand how distributed physical systems, such as the brain, can compute credit signals using only local coupling and local dynamics. PC-ALM may also inform deep learning in neuromorphic hardware, where dynamics are cheaper than on GPUs. Paper: https://arxiv.org/abs/2605.31022 Code: https://github.com/SakanaAI/pc-alm
    @nikparth1If I had a nickel for every NeuroAI method that "works" when training some network on MNIST and then we never see a follow-up getting it to work on any real problem.....
    @kaixhinJeffrey has been on a mission, investigating and developing local learning rules. We know that the brain can't do backprop - at least not exactly - so it would be amazing if we could figure out what it is doing instead.
    @mayferquick MNIST test
    @PMinervini@SakanaAILabs this idea has been around for a few years tho no? e.g. https://arxiv.org/abs/1808.06934
    @pfauRT @jeffreyseely: sharing a new paper! Augmented Lagrangian Predictive Coding I think one of the coolest unsolved problems is how the bra…
    @YesThisIsLionRT @kaixhin: Jeffrey has been on a mission, investigating and developing local learning rules. We know that the brain can't do backprop - a…
    @ylecunNice! This ends up being a version of what some of us have called "target prop": every layer's input is a free latent variable that serves as a target for the previous layer. As this paper points out, this can be derived from an "augmented Lagrangian" formulation of backprop in which the constraints (input of layer k+1 = output of layer k) are turned into penalties (divergence between input of layer k+1 and output of layer k). I've always hoped more people would pick up on this idea. I'm happy this is happening! I must say though that target prop, in the end, optimizes the same criterion as backprop and does the same thing as backprop while evaluating the gradient in a different way, perhaps more biologically plausible. My lab did some work on this idea in the context of "sparse auto-encoders" in the late 2000s. It turns out when the code in an auto-encoder is regularized (e.g. with L1 to make it sparse) target prop seems more efficient than backprop. https://scholar.google.com/citations?view_op=view_citation&hl=en&user=WLN3QrAAAAAJ&cstart=400&pagesize=100&sortby=pubdate&citation_for_view=WLN3QrAAAAAJ:JQOojiI6XY0C
    @Stefania_drugaGreat validation of my colleagues work @jeffreyseely & Julian Gould @SakanaAILabs
    @aran_nayebiRT @nikparth1: If I had a nickel for every NeuroAI method that "works" when training some network on MNIST and then we never see a follow-u…

    16 Sources

    @SakanaAILabsIntroducing PC-ALM, a local-learning alternative to backpropagation. Our method trains 1000-layer neural nets using only local dynamics, and without backprop. Blog: https://pub.sakana.ai/pc-alm/ Standard deep learning relies on backpropagation. The brain, however, cannot implement backpropagation, at least not exactly. How can a physical system, such as the brain, solve multilayer credit assignment without explicit use of backprop? We look for inspiration in two related fields: distributed optimization and NeuroAI. In NeuroAI, predictive coding asks each neuron activation to solve an energy-based inference problem instead of using a standard forward pass. That inference step can be implemented as energy-minimization dynamics on local prediction errors. This perspective -- each layer as a dynamical system -- has proven promising, but performance of predictive coding hasn't scaled well with depth. Credit signals at far ends of the network struggle to diffuse into internal layers. We turn to distributed optimization, generalizing predictive coding to use an augmented Lagrangian instead of energy. This motivation stems back to a classic 1988 paper by LeCun, showing that the Lagrange multipliers of a deep network can be identified with gradients of a supervised loss. The augmented Lagrangian then bridges LeCun's perspective to the standard predictive coding that is used in NeuroAI. We find that this new perspective yields a natural PC-like alternative to backpropagation, resulting in a method we call PC-ALM. PC-ALM differs from PC in that it introduces dual neurons (Lagrange multipliers) as part of the layer-local dynamics, resulting in each layer acting as a PI feedback control system to minimize local prediction errors. We find that PC-ALM is capable of propagating signals to seemingly arbitrary depth, especially in deep narrow networks where standard PC struggles to learn. Ultimately, our motivation here is to understand how distributed physical systems, such as the brain, can compute credit signals using only local coupling and local dynamics. PC-ALM may also inform deep learning in neuromorphic hardware, where dynamics are cheaper than on GPUs. Paper: https://arxiv.org/abs/2605.31022 Code: https://github.com/SakanaAI/pc-alm
    @nikparth1If I had a nickel for every NeuroAI method that "works" when training some network on MNIST and then we never see a follow-up getting it to work on any real problem.....
    @kaixhinJeffrey has been on a mission, investigating and developing local learning rules. We know that the brain can't do backprop - at least not exactly - so it would be amazing if we could figure out what it is doing instead.
    @mayferquick MNIST test
    @PMinervini@SakanaAILabs this idea has been around for a few years tho no? e.g. https://arxiv.org/abs/1808.06934
    @pfauRT @jeffreyseely: sharing a new paper! Augmented Lagrangian Predictive Coding I think one of the coolest unsolved problems is how the bra…
    @YesThisIsLionRT @kaixhin: Jeffrey has been on a mission, investigating and developing local learning rules. We know that the brain can't do backprop - a…
    @ylecunNice! This ends up being a version of what some of us have called "target prop": every layer's input is a free latent variable that serves as a target for the previous layer. As this paper points out, this can be derived from an "augmented Lagrangian" formulation of backprop in which the constraints (input of layer k+1 = output of layer k) are turned into penalties (divergence between input of layer k+1 and output of layer k). I've always hoped more people would pick up on this idea. I'm happy this is happening! I must say though that target prop, in the end, optimizes the same criterion as backprop and does the same thing as backprop while evaluating the gradient in a different way, perhaps more biologically plausible. My lab did some work on this idea in the context of "sparse auto-encoders" in the late 2000s. It turns out when the code in an auto-encoder is regularized (e.g. with L1 to make it sparse) target prop seems more efficient than backprop. https://scholar.google.com/citations?view_op=view_citation&hl=en&user=WLN3QrAAAAAJ&cstart=400&pagesize=100&sortby=pubdate&citation_for_view=WLN3QrAAAAAJ:JQOojiI6XY0C
    @Stefania_drugaGreat validation of my colleagues work @jeffreyseely & Julian Gould @SakanaAILabs
    @aran_nayebiRT @nikparth1: If I had a nickel for every NeuroAI method that "works" when training some network on MNIST and then we never see a follow-u…