Negative Self-Distillation aims to teach LLMs what not to do
A post introduces Negative Self-Distillation (NSD) as a label-free method for training large language models to explicitly avoid flaws.
TLDR
The announcement asks whether LLMs could learn to reason by learning what not to do, rather than copying perfect solutions. It presents Negative Self-Distillation (NSD) as a label-free training method focused on avoiding flaws.
Combined views
16.9K
1 Source, first seen 16d ago
140 likes