Claude Fable 5.1 more than doubled its predecessor’s benchmark score, The New Stack reports
On a modest budget, the model passed one of five tasks—but its failures cost less, The New Stack says.
TLDR
The New Stack reports that Claude Fable 5.1 more than doubled its predecessor’s benchmark score. On a modest budget, it passed one of five tasks, though the publication says its failures cost less.
Combined views
1K
2 Sources, first seen 21d ago
1 likes