Claude Fable 5.1 Tops Senior SWE-Bench Leaderboard
Post links medium effort approach to new model benchmark results.
TLDR
Dan Cleary of Anthropic quoted a post by Henry Ehrenberg of SnorkelAI on using medium effort as a starting point. The attached generated summary states that Claude Fable 5.1 reached first place on the Senior SWE-Bench leaderboard by raising consistency on the pass^3 metric versus Fable 5. The same summary adds that the model also leads the EMB finance benchmark with gains in Excel modeling.
Combined views
360
1 Source, first seen 28d ago