Claude reportedly passes fewer than a quarter of tests on an agent-building benchmark
The New Stack reports that Claude nevertheless performed best on Sierra’s Hyper-τ-bench, which tests how well AI agents can build other agents.
TLDR
Sierra has open-sourced Hyper-τ-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents, The New Stack reports. The outlet says Claude performed best on the new benchmark but passed fewer than a quarter of its tests.
Combined views
405
1 Source, first seen 5h ago
Claude reportedly passes fewer than a quarter of tests on an agent-building benchmark
The New Stack reports that Claude nevertheless performed best on Sierra’s Hyper-τ-bench, which tests how well AI agents can build other agents.
TLDR
Sierra has open-sourced Hyper-τ-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents, The New Stack reports. The outlet says Claude performed best on the new benchmark but passed fewer than a quarter of its tests.