• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
AI

Top coding-agent setup reportedly passes 23.9% of ττ-bench’s held-out customer simulations

A post summarizing ττ-bench says Claude Opus 5 with Claude Code passed 23.9% of held-out customer simulations across 53 tasks—the best result tested, versus 82.2% for the expert-authored reference.

3 Sources, 19d ago, first seen 19d ago

TLDR

A post describing ττ-bench says the benchmark asks coding agents to build an agent using scattered company records, a client, an API, existing code and a budget, then handle unseen customer requests. It reports that the top setup, Claude Opus 5 with Claude Code, passed 23.9% of held-out customer simulations across 53 tasks, compared with 82.2% for the expert-authored reference. The author argues that better code generation alone is not enough, pointing to agents barely questioning clients, reusing familiar designs and writing tests that missed their own errors.

Combined views

—

3 Sources, first seen 19d ago

— likes— comments— saves— reposts

Combined views

—

3 Sources, first seen 19d ago

— likes— comments— saves— reposts

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

Sentiment

Positive——Negative

Summary

Not enough discussion yet.

No sentiment analysis available yet.

3 Sources

Rohan Paul@rohanpaul_aiThe best coding-agent setup passed only 23.9% of held-out customer simulations on this benchmark. ττ-bench finds that today's coding agents fail most realistic agent-building work because they do not understand requirements deeply enough, so improving code generation alone will not solve the problem. ττ-bench treats agent building like a real client job. The model gets scattered company records, a client, an API, existing code, and a budget, then must ship an agent for unseen customer requests. The best setup, Claude Opus 5 with Claude Code, passed just 23.9% of held-out customer simulations across 53 tasks, versus 82.2% for the expert-authored reference. Most failures were outside pure coding. Agents searched records instead of understanding them, barely questioned the client, reused familiar designs, and wrote tests that often missed their own errors. On client-enabled tasks, builds asking 0 questions averaged 0.16; those asking 4+ averaged 0.50. So, better coding models are not enough, we also need coding agents that gather requirements, compare designs, use budgets intelligently, and run tests that expose their own blind spots before deployment.19d
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI

    3 Sources

    Rohan Paul@rohanpaul_aiThe best coding-agent setup passed only 23.9% of held-out customer simulations on this benchmark. ττ-bench finds that today's coding agents fail most realistic agent-building work because they do not understand requirements deeply enough, so improving code generation alone will not solve the problem. ττ-bench treats agent building like a real client job. The model gets scattered company records, a client, an API, existing code, and a budget, then must ship an agent for unseen customer requests. The best setup, Claude Opus 5 with Claude Code, passed just 23.9% of held-out customer simulations across 53 tasks, versus 82.2% for the expert-authored reference. Most failures were outside pure coding. Agents searched records instead of understanding them, barely questioned the client, reused familiar designs, and wrote tests that often missed their own errors. On client-enabled tasks, builds asking 0 questions averaged 0.16; those asking 4+ averaged 0.50. So, better coding models are not enough, we also need coding agents that gather requirements, compare designs, use budgets intelligently, and run tests that expose their own blind spots before deployment.19d
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet