Report
Mixing Claude Code and Codex reportedly raised fully correct patches from 45.8% to 62.5% at matched spend
A post describing a Meta paper says cross-review caught more silent bugs than giving one coding agent a bigger budget.
TLDR
A post describing a Meta paper says using Claude Code and Codex on the same task raised fully correct patches from 45.8% to 62.5% at matched spend. It attributes the gain to agents making different mistakes, giving cross-review more to catch, and says the effect repeated on an unrelated training codebase.
Combined views
3.3K
1 Source, first seen ago
