Model reportedly gains 5.7 percentage points on SWE-Bench Pro after office-work training
Surge AI says the model's additional training covered documents, spreadsheets, research, planning and tool use—with zero coding tasks. It attributes the coding benchmark gain to more general task-execution skills.
TLDR
Surge AI reports a 5.7-percentage-point improvement on SWE-Bench Pro after post-training a model on long-horizon office work, with no coding tasks. Based on the model's task trajectories, the company argues that it improved at general execution rather than learning more software engineering. Surge AI calls this “Goal-Directed Execution”: forming the right goals, building an accurate picture of the environment, keeping the main objective intact and checking that the job was actually done.
