Does a Claude model really need Claude Code?
Arena describes a comparison of Claude Code, Codex and Pi. The researchers say harness choice had little effect on task success rate but can significantly affect cost.
TLDR
Researchers tested seven models across Claude Code, Codex and Pi—three coding-agent harnesses, or software setups used to run the agents. Arena says the evaluation covered 21 model-harness pairs. The researchers report that harness choice had little effect on task success rate but can significantly affect cost. They also say a simple harness can be competitive, and a model’s native harness isn’t always best.
Combined views
597.2K
20 Sources, first seen ago