• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Does a Claude model really need Claude Code?

    Arena describes a comparison of Claude Code, Codex and Pi. The researchers say harness choice had little effect on task success rate but can significantly affect cost.

    AKAK
    Omar KhattabOK
    Arena.aiAR
    20 Sources, ,

    TLDR

    Researchers tested seven models across Claude Code, Codex and Pi—three coding-agent harnesses, or software setups used to run the agents. Arena says the evaluation covered 21 model-harness pairs. The researchers report that harness choice had little effect on task success rate but can significantly affect cost. They also say a simple harness can be competitive, and a model’s native harness isn’t always best.

    Combined views

    597.2K

    20 Sources, first seen 21d ago

    Combined views

    597.2K

    20 Sources, first seen 21d ago

    3.9K likes
    21d ago
    first seen 21d ago
    3.9K likes
    361 comments
    2.7K saves
    534 reposts
    361 comments
    2.7K saves
    534 reposts

    Sentiment

    Positive52%48%Negative

    Summary

    Positive accounts praised the model harness study for showing simple options can match native ones with fewer costs, while negative replies questioned the success metrics and evaluation of actual code output.

    Based on 27 sentiment-bearing replies from 21 accounts across 3 conversations.

    Sentiment

    Positive52%48%Negative

    Summary

    Positive accounts praised the model harness study for showing simple options can match native ones with fewer costs, while negative replies questioned the success metrics and evaluation of actual code output.

    Based on 27 sentiment-bearing replies from 21 accounts across 3 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    20 Sources

    Melissa Pan@melissapanDoes your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵21d
    Matei Zaharia@matei_zahariaRT @melissapan: Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising fi…21d
    Arena.ai@arenaDo models need their native harness for coding? Awesome work from our intern @melissapan on this. She dug into whether the harness (Claude Code vs Codex CLI vs Pi) actually moves the needle for coding agents. Turns out it matters way less than people assume. 21 model-harness pairs, real rigor. More from Arena to come. More details on the Arena blog: https://arena.ai/blog/coding-agents-harness-tax21d
    Ion Stoica@istoica05Does a model need its native harness? We evaluated seven models across the Claude Code, Codex, and Pi harnesses and found a surprising result: harness choice had little effect on task success but a substantial effect on cost. Sometimes, a simple harness is all you need!21d
    Andy Konwinski@andykonwinski"Harness choice has little effect on task success rate ... As models become more capable, agents may need less scaffolding."21d
    alex zhang@a1zhangRT @melissapan: Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising fi…20d
    Hamel Husain@HamelHusainRT @melissapan: Where is the harness tax coming from? 🧐 Agents can take similar numbers of turns at substantially different costs. For Fa…20d
    Tomasz Tunguz@ttunguzThe better the harness, the better the business. Berkeley published a study this week showing harnesses, the systems that control AI agents, set the price of an answer. The right harness cuts the cost of the same result by 71% without a loss of accuracy. https://x.com/i/article/210065185128068710420d
    Pi@pidotdev"Pi reaches the Pareto frontier on both benchmarks by providing just four tools: read, write, edit, and bash. To understand how harness design affects spending, we examine costs across completed attempts, recorded turn counts, and initial context." Sharing this analysis on the Harness Tax by @MelissaPan and others. https://harnesstax.github.io19d
    AK@_akhaliqAn Empirical Study of Harness Design for Coding Agents paper: https://huggingface.co/papers/2609.2080419d

    20 Sources

    Melissa Pan@melissapanDoes your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising findings emerge: 1️⃣Harness choice has little effect on task success rate, but can significantly affect the cost 2️⃣A simple harness can be competitive 3️⃣The native harness isn’t always the best. Millions of people are using coding agents, but the impact of harness choice remains unclear. (1/n) More details in the thread. 🧵21d
    Matei Zaharia@matei_zahariaRT @melissapan: Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising fi…21d
    Arena.ai@arenaDo models need their native harness for coding? Awesome work from our intern @melissapan on this. She dug into whether the harness (Claude Code vs Codex CLI vs Pi) actually moves the needle for coding agents. Turns out it matters way less than people assume. 21 model-harness pairs, real rigor. More from Arena to come. More details on the Arena blog: https://arena.ai/blog/coding-agents-harness-tax21d
    Ion Stoica@istoica05Does a model need its native harness? We evaluated seven models across the Claude Code, Codex, and Pi harnesses and found a surprising result: harness choice had little effect on task success but a substantial effect on cost. Sometimes, a simple harness is all you need!21d
    Andy Konwinski@andykonwinski"Harness choice has little effect on task success rate ... As models become more capable, agents may need less scaffolding."21d
    alex zhang@a1zhangRT @melissapan: Does your Claude model really need Claude Code…? 🤔 We evaluate 7 models on Claude Code, Codex, and Pi. Three surprising fi…20d
    Hamel Husain@HamelHusainRT @melissapan: Where is the harness tax coming from? 🧐 Agents can take similar numbers of turns at substantially different costs. For Fa…20d
    Tomasz Tunguz@ttunguzThe better the harness, the better the business. Berkeley published a study this week showing harnesses, the systems that control AI agents, set the price of an answer. The right harness cuts the cost of the same result by 71% without a loss of accuracy. https://x.com/i/article/210065185128068710420d
    Pi@pidotdev"Pi reaches the Pareto frontier on both benchmarks by providing just four tools: read, write, edit, and bash. To understand how harness design affects spending, we examine costs across completed attempts, recorded turn counts, and initial context." Sharing this analysis on the Harness Tax by @MelissaPan and others. https://harnesstax.github.io19d
    AK@_akhaliqAn Empirical Study of Harness Design for Coding Agents paper: https://huggingface.co/papers/2609.2080419d