Claude reportedly passes fewer than a quarter of tests on an agent-building benchmark
The New Stack reports that Claude nevertheless performed best on Sierraβs Hyper-Ο-bench, which tests how well AI agents can build other agents.
TLDR
Sierra has open-sourced Hyper-Ο-bench, a follow-up to its 2024 Ο-bench that tests how well AI agents can build other agents, The New Stack reports. The outlet says Claude performed best on the new benchmark but passed fewer than a quarter of its tests.
Combined views
405
1 Source, first seen 6h ago
Claude reportedly passes fewer than a quarter of tests on an agent-building benchmark
The New Stack reports that Claude nevertheless performed best on Sierraβs Hyper-Ο-bench, which tests how well AI agents can build other agents.
TLDR
Sierra has open-sourced Hyper-Ο-bench, a follow-up to its 2024 Ο-bench that tests how well AI agents can build other agents, The New Stack reports. The outlet says Claude performed best on the new benchmark but passed fewer than a quarter of its tests.