• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    “Antidoom” training aims to reduce small AI models’ “doom loops”

    A user says Liquid AI uses final token preference optimization (FTPO) to address loops that can occur when small models with thinking capabilities face complex tasks.

    LE
    1 Source, 15d ago, first seen 15d ago

    TLDR

    A user says small models with thinking capabilities can get stuck in “doom loops” on complex tasks, and that Liquid AI uses final token preference optimization (FTPO) as “antidoom” training. They share a tutorial described as explaining these loops, how to reduce them with FTPO, and how to implement a custom DPOTrainer class in TRL for the technique.

    Combined views

    16.4K

    1 Source, first seen 15d ago

    Combined views

    16.4K

    1 Source, first seen 15d ago

    377 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    377 likes
    17 comments
    346 saves
    45 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    17 comments
    346 saves
    45 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @helloiamleoniemodels can get stuck in "doom loops" when • the model is small, • it has thinking capabilities, • and the task is complex to overcome this issue, at @liquidai we use "antidoom" training using a technique called final token preference optimization (ftpo). in this tutorial, you will learn: • what is a doom loop • how to reduce doom loops with ftpo • how to implement a custom DPOTrainer class TRL for ftpo tutorial: https://github.com/Liquid4All/cookbook/blob/main/finetuning/notebooks/antidoom_with_trl.ipynb

    1 Source

    @helloiamleoniemodels can get stuck in "doom loops" when • the model is small, • it has thinking capabilities, • and the task is complex to overcome this issue, at @liquidai we use "antidoom" training using a technique called final token preference optimization (ftpo). in this tutorial, you will learn: • what is a doom loop • how to reduce doom loops with ftpo • how to implement a custom DPOTrainer class TRL for ftpo tutorial: https://github.com/Liquid4All/cookbook/blob/main/finetuning/notebooks/antidoom_with_trl.ipynb