Harness Choice Alters AI Coding Agent Rankings
Posts highlight study showing harness choice changes coding agent scores across models on SWE-bench Pro.
A retweet by Ofir Press shares a claim that Codex ranks second among ten harnesses with GLM-5.2 but ninth with Gemma-4. The accompanying source summary states a study tested ten coding agent harnesses on GLM-5.2 and Gemma-4 26B using SWE-bench Pro and found harness selection alters pass@1 scores. John Yang replies that his earlier work used simple ReAct setups and questions framing stronger harnesses as those with more tools. The visible posts focus on these benchmark observations without further confirmation of broader effects.
Combined views
13.8K
3 posts, first seen 23d ago