ProximalHQ Releases FrontierSWE v2 Long-Horizon Coding Benchmark
Benchmark tests models on extended autonomous software engineering tasks lasting up to 20 hours.
TLDR
ProximalHQ released FrontierSWE v2, an updated benchmark for ultra-long horizon coding tasks. The expanded suite uses improved methodology to evaluate frontier models on autonomous work up to 20 hours. Claude Fable 5.1 led other models by over 24 percentage points. The Proximal team reported insights on model performance at super-long tasks and introduced the Proximus harness based on mini-swe-agent, which delivered gains across models compared with standard harnesses.
Combined views
467.6K
25 Sources, first seen 28d ago