Manager-free coding-agent teams reportedly scored higher as they grew
A post describing a Microsoft paper says average scores rose at every step from one to 128 agents on the five hardest of ProgramBench's 200 tasks.
TLDR
A post describing a Microsoft paper says larger manager-free coding-agent teams scored higher and got there sooner. On ProgramBench's five hardest tasks, it says average scores rose at every step from one to 128 agents. Agensh agents claim subtasks, build and test work, merge it into a shared Git repo and log findings on a shared board. The post says each run used one model, and the paper did not report large-team costs.
Combined views
3.3K
1 Source, first seen ago
