Claude's agent-building benchmark lead despite passing fewer than a quarter of tests
The New Stack says Sierra has open-sourced Hyper-τ-bench, which tests how well AI agents can build other agents.
TLDR
The New Stack reports that Claude performed best on Hyper-τ-bench, a new benchmark for AI agents that build other agents, but passed fewer than a quarter of the tests. The outlet describes Sierra’s open-source benchmark as a follow-up to its 2024 τ-bench.
Combined views
2.2K
3 Sources, first seen ago
4 likes2 comments