• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Langford Calls Possible Overlap on Transformer Idea Parallel Development

    John Langford addresses rumor linking OpenAI model to full-bandwidth transformer work.

    SA
    JL
    RP
    7 Sources, 27d ago, first seen 27d ago

    TLDR

    John Langford stated that a rumor about GPT-6 Astra using a full-bandwidth transformer would represent parallel development if accurate. He noted his NeurIPS talk had already flagged recursion as a planned next step for investigation. Langford added he was asked about the rumor by someone from OpenAI and found the timing odd given his own prior agenda item. OpenAI separately announced GPT-6 Astra as its latest model focused on computer use, coding, cybersecurity, and science. An arXiv paper describes the full-bandwidth transformer approach, but no confirmation of its use in the model appears in the packet.

    Combined views

    135.8K

    7 Sources, first seen 27d ago

    Combined views

    135.8K

    7 Sources, first seen 27d ago

    908 likes
    908 likes
    35 comments
    967 saves
    127 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    35 comments
    967 saves
    127 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    7 Sources

    @gklambauerRECURRENT DEPTH IN GPT-6 ASTRA?? Transformer blocks with shared weights are applied several times. A so-called recurrent reasoning models (RRM), as we suggested earlier this year: P: https://arxiv.org/abs/2603.02193
    @JohnCLangfordYesterday, I was asked about a rumor that Astra https://openai.com/index/gpt-6-astra/ uses a Full Bandwidth Transformer https://arxiv.org/abs/2608.08888 . I have no special knowledge, but it is at least strange to have someone from OpenAI asking where the idea came from.
    @prfsanjeevaroraNew openAI architecture seems to have several close precedents in recent works including from our group https://arxiv.org/html/2607.15178v2
    @rohanpaul_aiSapient introduced the "Hierarchical Reasoning Model" in June 2025, then open-sourced its "Hierarchical Reasoning Model" Text model this May. The model uses recurrent architectural loops, while OpenAI GPT-6 Astra is now putting looped recurrence in latent space into the frontier conversation. The overlapping idea was: reasoning can spend computation by repeatedly updating internal state rather than relying only on a conventional transformer stack. Sapient made that research decision before it had much industry validation, then published a model people could actually inspect and retrain.

    7 Sources

    @gklambauerRECURRENT DEPTH IN GPT-6 ASTRA?? Transformer blocks with shared weights are applied several times. A so-called recurrent reasoning models (RRM), as we suggested earlier this year: P: https://arxiv.org/abs/2603.02193
    @JohnCLangfordYesterday, I was asked about a rumor that Astra https://openai.com/index/gpt-6-astra/ uses a Full Bandwidth Transformer https://arxiv.org/abs/2608.08888 . I have no special knowledge, but it is at least strange to have someone from OpenAI asking where the idea came from.
    @prfsanjeevaroraNew openAI architecture seems to have several close precedents in recent works including from our group https://arxiv.org/html/2607.15178v2
    @rohanpaul_aiSapient introduced the "Hierarchical Reasoning Model" in June 2025, then open-sourced its "Hierarchical Reasoning Model" Text model this May. The model uses recurrent architectural loops, while OpenAI GPT-6 Astra is now putting looped recurrence in latent space into the frontier conversation. The overlapping idea was: reasoning can spend computation by repeatedly updating internal state rather than relying only on a conventional transformer stack. Sapient made that research decision before it had much industry validation, then published a model people could actually inspect and retrain.
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet