‘Antidoom’ training aims to reduce small AI models’ ‘doom loops’
A Liquid AI team member says the company uses final token preference optimization (FTPO) to tackle loops that small models with thinking capabilities can enter on complex tasks.
TLDR
A Liquid AI team member shares a tutorial on “antidoom” training, saying the company uses final token preference optimization (FTPO) to address “doom loops” in small models with thinking capabilities. They describe the tutorial as covering what doom loops are, how to reduce them with FTPO, and how to implement a custom DPOTrainer class in TRL.
Combined views
17
2 Sources, first seen 15d ago