• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Extropic uses reinforcement learning to post-train models for Thermo AI research

    Extropic says the work aims to speed algorithmic discovery for what it calls a new species of computer.

    B(
    GV
    PI
    12 Sources, ,

    TLDR

    Extropic says it is using reinforcement learning to post-train models for Thermo AI research, aiming to accelerate algorithmic discovery for what it calls a “new species of computer.” It describes the work as the “first sparks of Thermo RSI,” in collaboration with Prime Intellect.

    Combined views

    215.6K

    12 Sources, first seen 3h ago

    Combined views

    215.6K

    12 Sources, first seen 3h ago

    1.9K likes
    3h ago
    first seen 3h ago
    1.9K likes
    132 comments
    437 saves
    206 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    132 comments
    437 saves
    206 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    #8

    Today's Rank

    #8

    12 Sources

    @extropicWhat evolves together fits together. We’re using RL to post-train models for Thermo AI research, accelerating algorithmic discovery for our new species of computer. The first sparks of Thermo RSI, in collaboration with @PrimeIntellect. https://extropic.ai/writing/baby-thermo-rsi
    @beffjezosRT @extropic: What evolves together fits together. We’re using RL to post-train models for Thermo AI research, accelerating algorithmic di…
    @PrimeIntellectUsing Prime Intellect, Extropic post-trained Qwen3.6-35B-A3B for thermodynamic ML research, nearly tripling its eval results on held-out tasks in ~100 GRPO steps. They built a custom RL environment with verifiers and trained on Hosted Training, Prime Sandboxes, and Prime Inference. This meant Extropic didn't have to manage multi-node GPU infrastructure, so the team could focus on research.
    @GillVerdArtificial Thermo Superintelligence assembling itself from the future. I love it when a temporal pincer comes together.

    12 Sources

    @extropicWhat evolves together fits together. We’re using RL to post-train models for Thermo AI research, accelerating algorithmic discovery for our new species of computer. The first sparks of Thermo RSI, in collaboration with @PrimeIntellect. https://extropic.ai/writing/baby-thermo-rsi
    @beffjezosRT @extropic: What evolves together fits together. We’re using RL to post-train models for Thermo AI research, accelerating algorithmic di…
    @PrimeIntellectUsing Prime Intellect, Extropic post-trained Qwen3.6-35B-A3B for thermodynamic ML research, nearly tripling its eval results on held-out tasks in ~100 GRPO steps. They built a custom RL environment with verifiers and trained on Hosted Training, Prime Sandboxes, and Prime Inference. This meant Extropic didn't have to manage multi-node GPU infrastructure, so the team could focus on research.
    @GillVerdArtificial Thermo Superintelligence assembling itself from the future. I love it when a temporal pincer comes together.