Announcement
Automatic hint optimization proposed as a step toward continual learning
Applied Compute says it optimizes teacher-model hints using efficient proxies for their eventual training value.
TLDR
Applied Compute describes a way to automatically optimize the hints a teacher model uses to supervise training in on-policy self-distillation. It says hint quality heavily affects the training signal, but refining hints manually across runs is impractical and isolating each hint’s impact is difficult. The company presents its method as a step toward continual learning.
Combined views
5.9K
2 Sources, first seen ago
