Claude Fable 5.1 more than doubled its predecessor’s benchmark score, The New Stack reports
On a modest budget, the model passed one of five tasks—but its failures cost less, The New Stack says.
On a modest budget, the model passed one of five tasks—but its failures cost less, The New Stack says.
The New Stack reports that Claude Fable 5.1 more than doubled its predecessor’s benchmark score. On a modest budget, it passed one of five tasks, though the publication says its failures cost less.
1K
2 posts, first seen 4d ago
On a modest budget, the model passed one of five tasks—but its failures cost less, The New Stack says.
The New Stack reports that Claude Fable 5.1 more than doubled its predecessor’s benchmark score. On a modest budget, it passed one of five tasks, though the publication says its failures cost less.
Not enough discussion yet.
No sentiment analysis available yet.
Not enough discussion yet.
No sentiment analysis available yet.
—
Not ranked yet
—
Not ranked yet