Negative Self-Distillation aims to train LLMs to avoid flaws
An introductory post describes NSD as a label-free method for training large language models to explicitly avoid flaws.
TLDR
A post introducing Negative Self-Distillation (NSD) asks whether large language models could learn to reason by learning what not to do, rather than copying perfect solutions. It presents NSD as a label-free training method built around avoiding flaws.
Combined views
1
1 Source, first seen 15d ago
reposts