Replit Agent's reported gains over a single-worker setup on coding benchmarks
Replit says its Agent lets the core model choose and reuse subagents as work unfolds. In its tests, it reports 11- and 16-point gains over a single-worker setup on DeepSWE and Terminal-Bench, respectively.
TLDR
Replit says its Agent lets GPT-6 Astra choose which subagents to use and how much effort to spend, rather than assigning it one long-lived worker. Replit reports scores of 72% on DeepSWE v1.1 and 49% on Terminal-Bench 4.0, beating its single-worker comparison by 11 and 16 points. It says Astra alone scored higher at its best settings, but at more than twice the cost per task.
Combined views
97.2K
2 Sources, first seen 3h ago