• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Amir Efrati Links OpenAI RLSlow to Q* Model

    The Information editor ties recent OpenAI mention to 2023 internal project.

    AE
    TB
    2 Sources, 24d ago, first seen 24d ago

    TLDR

    Amir Efrati posted that an OpenAI reference to RLSlow appears related to the Q* model created in 2023 by Jakub, Ilya, and Szymon. He noted that such internal breakthroughs generated both excitement and concern. Efrati added that loop transformers represent a more recent example. The post includes a mobile screenshot of OpenAI's September material and links to the original thread.

    Combined views

    49.2K

    2 Sources, first seen 24d ago

    Combined views

    49.2K

    2 Sources, first seen 24d ago

    485 likes
    485 likes
    12 comments
    225 saves
    32 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    12 comments
    225 saves
    32 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @amirThis OpenAI RLSlow reference appears to be related to the Q* model that Jakub Ilya and Szymon made in 2023, which we broke news on during l’affaire Altman a lot of internal breakthroughs cause both excitement and concern, including loop transformers more recently
    @TrapitBansalFun fact: RLSlow was named after Thinking, Fast and Slow. The idea was that language models already had a kind of "fast" thinking, producing an answer immediately, and that we could use RL to teach them "slow" thinking: deliberate, multi-token reasoning that spends more compute working through a problem. What I remember most from those days is how early the team developed real conviction in the direction, and how much work went into earning it. We were developing the algorithms, designing careful experiments to test the ideas, and watching the empirical evidence accumulate. These were many long nights babysitting runs, understanding what the results were telling us, and figuring out what to try next. A lot of those early discussions were with Ilya, and later with Jakub. Pretty early, the evidence had already pushed us to a strong view: RL for reasoning would scale. Models would learn to spend more compute at inference time to reason through increasingly hard problems, and this would fundamentally change how we think about inference. Many of us also spent countless hours reading through reasoning traces. Seeing how models arrived at answers, not just the answers themselves, felt like a powerful new lens on generalization and alignment. These ideas feel obvious now. They really weren’t then.

    2 Sources

    @amirThis OpenAI RLSlow reference appears to be related to the Q* model that Jakub Ilya and Szymon made in 2023, which we broke news on during l’affaire Altman a lot of internal breakthroughs cause both excitement and concern, including loop transformers more recently
    @TrapitBansalFun fact: RLSlow was named after Thinking, Fast and Slow. The idea was that language models already had a kind of "fast" thinking, producing an answer immediately, and that we could use RL to teach them "slow" thinking: deliberate, multi-token reasoning that spends more compute working through a problem. What I remember most from those days is how early the team developed real conviction in the direction, and how much work went into earning it. We were developing the algorithms, designing careful experiments to test the ideas, and watching the empirical evidence accumulate. These were many long nights babysitting runs, understanding what the results were telling us, and figuring out what to try next. A lot of those early discussions were with Ilya, and later with Jakub. Pretty early, the evidence had already pushed us to a strong view: RL for reasoning would scale. Models would learn to spend more compute at inference time to reason through increasingly hard problems, and this would fundamentally change how we think about inference. Many of us also spent countless hours reading through reasoning traces. Seeing how models arrived at answers, not just the answers themselves, felt like a powerful new lens on generalization and alignment. These ideas feel obvious now. They really weren’t then.