NoRA improves LoRA training with a one-line change, a paper summary says
A post describing Normalized Low-Rank Adaptation says the method adds no trainable parameters or extra computation when the model is used, while claiming more stable training and less forgetting.
TLDR
A post summarizing a paper on Normalized Low-Rank Adaptation (NoRA) says the method normalizes LoRA’s down-projection matrices during training. It reports faster convergence, better final performance, more stable training and less catastrophic forgetting—the loss of previously learned capabilities—across pretraining, supervised fine-tuning and reinforcement learning. The post also says applying the normalization just once, at initialization, improves standard LoRA and is cheaper than repeating it throughout training.
Combined views
10.4K
2 Sources, first seen 24d ago