Top coding-agent setup reportedly passes 23.9% of ττ-bench’s held-out customer simulations
A post summarizing ττ-bench says Claude Opus 5 with Claude Code passed 23.9% of held-out customer simulations across 53 tasks—the best result tested, versus 82.2% for the expert-authored reference.
TLDR
A post describing ττ-bench says the benchmark asks coding agents to build an agent using scattered company records, a client, an API, existing code and a budget, then handle unseen customer requests. It reports that the top setup, Claude Opus 5 with Claude Code, passed 23.9% of held-out customer simulations across 53 tasks, compared with 82.2% for the expert-authored reference. The author argues that better code generation alone is not enough, pointing to agents barely questioning clients, reusing familiar designs and writing tests that missed their own errors.
Combined views
4.3K
2 Sources, first seen 3h ago
Top coding-agent setup reportedly passes 23.9% of ττ-bench’s held-out customer simulations
A post summarizing ττ-bench says Claude Opus 5 with Claude Code passed 23.9% of held-out customer simulations across 53 tasks—the best result tested, versus 82.2% for the expert-authored reference.