• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Fernando Hernandez Garcia's PhD thesis tests ways to keep neural networks learning over time

    In the thesis abstract shared by Richard Sutton, Hernandez Garcia reports that neural networks lost learning ability across several architectures, including vision transformers.

    ('
    RS
    AC
    3 Sources, ,

    TLDR

    Richard Sutton announced that his student Fernando Hernandez Garcia's PhD thesis is available. The abstract describes plasticity loss, in which neural networks gradually lose their ability to learn from new data. Hernandez Garcia reports that selectively resetting parts of a network maintained its learning ability across the systems tested. Resetting units and resetting weights each had different trade-offs for stability and implementation.

    Combined views

    68.6K

    3 Sources, first seen 3h ago

    Combined views

    68.6K

    3 Sources, first seen 3h ago

    772 likes
    3h ago
    first seen 3h ago
    772 likes
    14 comments
    235 saves
    25 reposts
    14 comments
    235 saves
    25 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #6

    Today's Rank

    #6

    3 Sources

    @RichardSSuttonThe PhD thesis of my 16th PhD student, Fernando Hernandez Garcia, is now available. Title: Selective Reinitialization Algorithms for Preventing Plasticity Loss in Artificial Neural Networks Url: http://www.incompleteideas.net/papers/FernandoPhD.pdf Abstract: In this dissertation, I study systems based on artificial neural networks that learn from nonstationary data. Learning from non-stationary data requires continual adaptation of the system, and is often referred to as continual learning. Developing systems capable of continual learning is a longstanding goal in artificial intelligence. In deep learning, the field concerned with designing and training deep neural networks, this goal remains elusive. This dissertation addresses a fundamental challenge in continual learning: loss of plasticity, a phenomenon where artificial neural network systems progressively lose their ability to learn from new data. While the phenomenon has been noted several times over the last three decades, it has remained understudied until recently. The work in this dissertation constitutes the first systematic demonstrations of plasticity loss, highlighting its persistence and importance. In this work, I provide a systematic demonstration of plasticity loss in a wide variety of deep learning systems. Across systems based on fully-connected networks, convolutional networks, residual networks, and vision transformers, plasticity degrades when learning continually. Notably, even systems employing normalization techniques, residual connections, and regularization—design choices that improve training stability—remain susceptible to plasticity loss. This evidence establishes that plasticity loss is pervasive and that deep learning systems trained with backpropagation are not suitable for continual learning. I explore the idea of selective reinitialization to prevent plasticity loss. This idea has been used in the past for improving generalization performance in deep learning systems. However, its use for preventing plasticity loss is a recent innovation pioneered by the continual backpropagation algorithm. The algorithm periodically sets new values for units in the network, making initialization a continuous process rather than a one-time operation. I demonstrate the effectiveness of continual backpropagation in preventing plasticity loss across a wide variety of settings. These demonstrations establish that plasticity loss, while pervasive, is not inherent to deep learning systems. I generalize continual backpropagation through an algorithm I call selective unit reinitialization. This general algorithm has three key components: a utility measure that ranks units by importance, a pruning criterion that selects which units to reinitialize, and a reinitialization method that assigns new values to the selected units. Continual backpropagation involves a specific choice of utility measure, pruning criterion, and reinitialization method. However, the general algorithm is not limited to the choices used in continual backpropagation. I present a study of how different choices for the three components of selective unit reinitialization affect its effectiveness at maintaining plasticity. This study establishes selective unit reinitialization as a general approach that can be tailored to each learning system to maintain plasticity. Finally, I propose a different approach to implementing selective reinitialization that operates at the weight level. I call the corresponding algorithm selective weight reinitialization. Reinitializing at the unit or weight level involves different trade-offs. Unit reinitialization minimally disrupts the network outputs, preserving stability during learning. However, the definition of a unit varies across network architectures, requiring additional engineering to make the approach effective. Weight reinitialization, in contrast, can be readily applied to any arbitrary network architecture, but can substantially affect network outputs, reducing training stability. The reduced learning stability can be remedied with L2 regularization, which stabilizes selective weight reinitialization but introduces an additional hyperparameter. Both selective unit and weight reinitialization successfully maintained plasticity across the systems tested, providing flexible approaches for different architectural and engineering constraints. Fernando is now a research scientist at Zyphra.3h
    @yoavgowait what?? faculty in post-LLM AI are farting out 10 phds/year and Richard Sutton had 16 graduating students in his entire career?? this is amazing. i respect him so much more now3h
    @andrew_n_carr@yoavgo He's my academic great grandfather and I also respect him immensely1h

    3 Sources

    @RichardSSuttonThe PhD thesis of my 16th PhD student, Fernando Hernandez Garcia, is now available. Title: Selective Reinitialization Algorithms for Preventing Plasticity Loss in Artificial Neural Networks Url: http://www.incompleteideas.net/papers/FernandoPhD.pdf Abstract: In this dissertation, I study systems based on artificial neural networks that learn from nonstationary data. Learning from non-stationary data requires continual adaptation of the system, and is often referred to as continual learning. Developing systems capable of continual learning is a longstanding goal in artificial intelligence. In deep learning, the field concerned with designing and training deep neural networks, this goal remains elusive. This dissertation addresses a fundamental challenge in continual learning: loss of plasticity, a phenomenon where artificial neural network systems progressively lose their ability to learn from new data. While the phenomenon has been noted several times over the last three decades, it has remained understudied until recently. The work in this dissertation constitutes the first systematic demonstrations of plasticity loss, highlighting its persistence and importance. In this work, I provide a systematic demonstration of plasticity loss in a wide variety of deep learning systems. Across systems based on fully-connected networks, convolutional networks, residual networks, and vision transformers, plasticity degrades when learning continually. Notably, even systems employing normalization techniques, residual connections, and regularization—design choices that improve training stability—remain susceptible to plasticity loss. This evidence establishes that plasticity loss is pervasive and that deep learning systems trained with backpropagation are not suitable for continual learning. I explore the idea of selective reinitialization to prevent plasticity loss. This idea has been used in the past for improving generalization performance in deep learning systems. However, its use for preventing plasticity loss is a recent innovation pioneered by the continual backpropagation algorithm. The algorithm periodically sets new values for units in the network, making initialization a continuous process rather than a one-time operation. I demonstrate the effectiveness of continual backpropagation in preventing plasticity loss across a wide variety of settings. These demonstrations establish that plasticity loss, while pervasive, is not inherent to deep learning systems. I generalize continual backpropagation through an algorithm I call selective unit reinitialization. This general algorithm has three key components: a utility measure that ranks units by importance, a pruning criterion that selects which units to reinitialize, and a reinitialization method that assigns new values to the selected units. Continual backpropagation involves a specific choice of utility measure, pruning criterion, and reinitialization method. However, the general algorithm is not limited to the choices used in continual backpropagation. I present a study of how different choices for the three components of selective unit reinitialization affect its effectiveness at maintaining plasticity. This study establishes selective unit reinitialization as a general approach that can be tailored to each learning system to maintain plasticity. Finally, I propose a different approach to implementing selective reinitialization that operates at the weight level. I call the corresponding algorithm selective weight reinitialization. Reinitializing at the unit or weight level involves different trade-offs. Unit reinitialization minimally disrupts the network outputs, preserving stability during learning. However, the definition of a unit varies across network architectures, requiring additional engineering to make the approach effective. Weight reinitialization, in contrast, can be readily applied to any arbitrary network architecture, but can substantially affect network outputs, reducing training stability. The reduced learning stability can be remedied with L2 regularization, which stabilizes selective weight reinitialization but introduces an additional hyperparameter. Both selective unit and weight reinitialization successfully maintained plasticity across the systems tested, providing flexible approaches for different architectural and engineering constraints. Fernando is now a research scientist at Zyphra.3h
    @yoavgowait what?? faculty in post-LLM AI are farting out 10 phds/year and Richard Sutton had 16 graduating students in his entire career?? this is amazing. i respect him so much more now3h
    @andrew_n_carr@yoavgo He's my academic great grandfather and I also respect him immensely1h