• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    NoRA improves LoRA training with a one-line change, a paper summary says

    A post describing Normalized Low-Rank Adaptation says the method adds no trainable parameters or extra computation when the model is used, while claiming more stable training and less forgetting.

    EL
    2 Sources, 24d ago, first seen 24d ago

    TLDR

    A post summarizing a paper on Normalized Low-Rank Adaptation (NoRA) says the method normalizes LoRA’s down-projection matrices during training. It reports faster convergence, better final performance, more stable training and less catastrophic forgetting—the loss of previously learned capabilities—across pretraining, supervised fine-tuning and reinforcement learning. The post also says applying the normalization just once, at initialization, improves standard LoRA and is cheaper than repeating it throughout training.

    Combined views

    10.4K

    2 Sources, first seen 24d ago

    Combined views

    10.4K

    2 Sources, first seen 24d ago

    108 likes
    108 likes
    16 comments
    82 saves
    29 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    16 comments
    82 saves
    29 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @omarsar0// Normalized Low-Rank Adaptation (NoRA) // They propose a one-line change to LoRA that costs nothing and improves convergence, stability and forgetting. LoRA initializes the up-projection to zero, which means early optimization is governed almost entirely by the down-projection. This observation tells you where to regularize. NoRA normalizes the down-projection matrices during training. The authors also show the same normalization applied once at initialization improves standard LoRA without repeating it through training, which is the cheaper of the two options. The benefits hold across pretraining, supervised fine-tuning and reinforcement learning. Faster convergence, better final performance, more stable training, and less catastrophic forgetting. It adds no trainable parameters and no inference-time computation, which is what makes it broadly applicable rather than another specialized LoRA variant. Paper: https://academy.dair.ai/papers/normalized-low-rank-adaptation-2608.31036

    2 Sources

    @omarsar0// Normalized Low-Rank Adaptation (NoRA) // They propose a one-line change to LoRA that costs nothing and improves convergence, stability and forgetting. LoRA initializes the up-projection to zero, which means early optimization is governed almost entirely by the down-projection. This observation tells you where to regularize. NoRA normalizes the down-projection matrices during training. The authors also show the same normalization applied once at initialization improves standard LoRA without repeating it through training, which is the cheaper of the two options. The benefits hold across pretraining, supervised fine-tuning and reinforcement learning. Faster convergence, better final performance, more stable training, and less catastrophic forgetting. It adds no trainable parameters and no inference-time computation, which is what makes it broadly applicable rather than another specialized LoRA variant. Paper: https://academy.dair.ai/papers/normalized-low-rank-adaptation-2608.31036