• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    ProximalHQ Releases FrontierSWE v2 Long-Horizon Coding Benchmark

    Benchmark tests models on extended autonomous software engineering tasks lasting up to 20 hours.

    NI
    RP
    BC
    25 Sources, 28d ago, first seen 28d ago

    TLDR

    ProximalHQ released FrontierSWE v2, an updated benchmark for ultra-long horizon coding tasks. The expanded suite uses improved methodology to evaluate frontier models on autonomous work up to 20 hours. Claude Fable 5.1 led other models by over 24 percentage points. The Proximal team reported insights on model performance at super-long tasks and introduced the Proximus harness based on mini-swe-agent, which delivered gains across models compared with standard harnesses.

    Combined views

    467.6K

    25 Sources, first seen 28d ago

    Combined views

    467.6K

    25 Sources, first seen 28d ago

    3.3K likes
    3.3K likes
    183 comments
    822 saves
    237 reposts

    Sentiment

    Positive89.4%10.6%Negative

    Based on 50 sentiment-bearing replies from 47 accounts across 3 conversations.

    183 comments
    822 saves
    237 reposts

    Sentiment

    Positive89.4%10.6%Negative

    Based on 50 sentiment-bearing replies from 47 accounts across 3 conversations.

    25 Sources

    @ProximalHQWe are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin
    @calvinchenReally excited to launch FrontierSWE v2! In one task, we ask models to build an OpenGL engine that can render a flight simulator game. Excited to share more detailed analysis in the coming days but checkout the blog for more.
    @samsja19great work
    @nrehiew_Excited to finally release FrontierSWE v2. We evaluate models on extremely difficult, extremely long-horizon tasks where models work autonomously for up to 20 hours. We evaluated Fable 5.1 and found that it was the best available model by over 24 percentage points
    @aryaman2020really great benchmark, with interesting analysis of reward hacks that we don’t get to see in public often!
    @SuhailExcited to see benchmarks headed in this direction and getting better!
    @mobav0RT @ProximalHQ: We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and im…
    @khoomeikIs science just a category of SWE tasks? OAI/Ant hillclimbing on tasks like this will certainly accelerate scientific workflows. But the more I see scientists do their work at Periodic, the more I come to appreciate the difficulty of verification in science. While many of the enabling workflows & computational subtasks in a research project have clear verifiers, dealing with uncertainty in the physical world is what makes science hard. Unlike math, code, or simulation, the world of atoms does not have an oracle. If you want to hillclimb against it, you need deep scientific expertise, a high-throughput lab, and maybe even new ML methods. Congrats to @MatternJustus and the Proximal team on another great benchmark! I'm excited to see the community continue to accelerate the verifiable parts of science.
    @0interestratesRT @calvinchen: Really excited to launch FrontierSWE v2! In one task, we ask models to build an OpenGL engine that can render a flight sim…
    @xeophonreading traces is so fun (the prompt does not talk about littlefs, the model inferred it from the tests + pre-train knowledge)

    25 Sources

    @ProximalHQWe are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and improved methodology. We see large performance gaps between frontier models, with Claude Fable 5.1 leading by a wide margin
    @calvinchenReally excited to launch FrontierSWE v2! In one task, we ask models to build an OpenGL engine that can render a flight simulator game. Excited to share more detailed analysis in the coming days but checkout the blog for more.
    @samsja19great work
    @nrehiew_Excited to finally release FrontierSWE v2. We evaluate models on extremely difficult, extremely long-horizon tasks where models work autonomously for up to 20 hours. We evaluated Fable 5.1 and found that it was the best available model by over 24 percentage points
    @aryaman2020really great benchmark, with interesting analysis of reward hacks that we don’t get to see in public often!
    @SuhailExcited to see benchmarks headed in this direction and getting better!
    @mobav0RT @ProximalHQ: We are releasing FrontierSWE v2, our updated ultra-long horizon coding benchmark V2 features an expanded task suite and im…
    @khoomeikIs science just a category of SWE tasks? OAI/Ant hillclimbing on tasks like this will certainly accelerate scientific workflows. But the more I see scientists do their work at Periodic, the more I come to appreciate the difficulty of verification in science. While many of the enabling workflows & computational subtasks in a research project have clear verifiers, dealing with uncertainty in the physical world is what makes science hard. Unlike math, code, or simulation, the world of atoms does not have an oracle. If you want to hillclimb against it, you need deep scientific expertise, a high-throughput lab, and maybe even new ML methods. Congrats to @MatternJustus and the Proximal team on another great benchmark! I'm excited to see the community continue to accelerate the verifiable parts of science.
    @0interestratesRT @calvinchen: Really excited to launch FrontierSWE v2! In one task, we ask models to build an OpenGL engine that can render a flight sim…
    @xeophonreading traces is so fun (the prompt does not talk about littlefs, the model inferred it from the tests + pre-train knowledge)
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet