“Antidoom” training aims to reduce small AI models’ “doom loops”
A user says Liquid AI uses final token preference optimization (FTPO) to address loops that can occur when small models with thinking capabilities face complex tasks.
TLDR
A user says small models with thinking capabilities can get stuck in “doom loops” on complex tasks, and that Liquid AI uses final token preference optimization (FTPO) as “antidoom” training. They share a tutorial described as explaining these loops, how to reduce them with FTPO, and how to implement a custom DPOTrainer class in TRL for the technique.
Combined views
16.4K
1 Source, first seen 15d ago