• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Researcher Questions OpenAI Reasoning Methods

    Researcher at Pangram Labs posts skepticism on OpenAI model techniques.

    ZE
    UE
    2 Sources, 26d ago, first seen 26d ago

    TLDR

    ueaj, a researcher at Pangram Labs, tweeted that OpenAI is probably not using looped transformers. The post allows it might involve something like MoEUT but states that approach would not account for denser reasoning observed in frontier math tasks. ueaj added that the RL signal appears too sparse to achieve high performance with minimal reasoning steps. The message was retweeted by commentator zephyr_z9. The post includes a diagram attachment but provides no further confirmation of any OpenAI architecture.

    Combined views

    31.4K

    2 Sources, first seen 26d ago

    Combined views

    31.4K

    2 Sources, first seen 26d ago

    433 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    433 likes
    17 comments
    342 saves
    49 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    17 comments
    342 saves
    49 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    2 Sources

    @_ueajI don't think OpenAI is doing looped transformers, though it is possible they're doing something like MoEUT. However, that doesn't really explain how they are actually eliciting significantly denser reasoning. I don't think the RL signal is dense enough to saturate frontier math with minimal reasoning or perform complex actions without reasoning at all. Instead, it's far more likely they're doing synthetic rewriting of CoTs. Rewriting them to be terser, or outright dropping parts of the CoT and asking the next generation pretrain to fill in the middle, essentially internalizing the reasoning patterns deep in the weights. This is likely to me for many reasons, the first is that it gets around the RL signal problem. You turn a RL task into a dense SFT/pretraining task, so the signal is far denser and you get more bits per episode. The second is that it's simple and true breakthroughs are rare, and it's something that needs to happen at scale (you need pretraining scale data for this to work). If you're markov brained, the original CoT is our data generating process, and by dropping parts of it, we force the model to reconstruct the hidden variables (the hidden parts of the CoT) but in the model weights instead. (blog below)
    @zephyr_z9RT @_ueaj: I don't think OpenAI is doing looped transformers, though it is possible they're doing something like MoEUT. However, that doesn…

    2 Sources

    @_ueajI don't think OpenAI is doing looped transformers, though it is possible they're doing something like MoEUT. However, that doesn't really explain how they are actually eliciting significantly denser reasoning. I don't think the RL signal is dense enough to saturate frontier math with minimal reasoning or perform complex actions without reasoning at all. Instead, it's far more likely they're doing synthetic rewriting of CoTs. Rewriting them to be terser, or outright dropping parts of the CoT and asking the next generation pretrain to fill in the middle, essentially internalizing the reasoning patterns deep in the weights. This is likely to me for many reasons, the first is that it gets around the RL signal problem. You turn a RL task into a dense SFT/pretraining task, so the signal is far denser and you get more bits per episode. The second is that it's simple and true breakthroughs are rare, and it's something that needs to happen at scale (you need pretraining scale data for this to work). If you're markov brained, the original CoT is our data generating process, and by dropping parts of it, we force the model to reconstruct the hidden variables (the hidden parts of the CoT) but in the model weights instead. (blog below)
    @zephyr_z9RT @_ueaj: I don't think OpenAI is doing looped transformers, though it is possible they're doing something like MoEUT. However, that doesn…