• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
Technology

Claude reportedly passes fewer than a quarter of tests on an agent-building benchmark

The New Stack reports that Claude nevertheless performed best on Sierra’s Hyper-τ-bench, which tests how well AI agents can build other agents.

The New StackTN
1 Source, 20d ago, first seen 20d ago

TLDR

Sierra has open-sourced Hyper-τ-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents, The New Stack reports. The outlet says Claude performed best on the new benchmark but passed fewer than a quarter of its tests.

Combined views

639

1 Source, first seen 20d ago

Combined views

639

1 Source, first seen 20d ago

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

1 Source

The New Stack@thenewstackClaude performed best on a new benchmark for 'agents that build agents'. But it passed fewer than a quarter of the tests. https://thenewstack.io/claude-build-agents-benchmark/?taid=6aac70f8c1df6d0001468c8e&utm_campaign=trueanthem&utm_medium=social&utm_source=twitter20d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    1 Source

    The New Stack@thenewstackClaude performed best on a new benchmark for 'agents that build agents'. But it passed fewer than a quarter of the tests. https://thenewstack.io/claude-build-agents-benchmark/?taid=6aac70f8c1df6d0001468c8e&utm_campaign=trueanthem&utm_medium=social&utm_source=twitter20d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet