Reaction
Off-policy training is claimed to work well for continual AI model learning
A post says supervised fine-tuning on heavily post-trained models often leads to forgetting and poor generalization.
TLDR
The post says many approaches try to make training data more on-policy because heavily post-trained models often forget and fail to generalize under supervised fine-tuning. It argues that balancing on-policy data with quality is tricky, and claims off-policy training can work well for continual learning if done right.
Combined views
4.7K
2 Sources, first seen ago