CUA-Speedrun aims to standardize speed tests for computer-use agents
Its creators say the toolkit fixes hardware, desktops, tasks and timing to make agent runtimes easier to compare.
TLDR
CUA-Speedrun’s creators say evaluation setups can heavily affect measured runtime even when agents achieve similar task-success scores. Their toolkit uses fixed hardware, desktops, tasks and timing across four computer-use benchmarks. They also report that a faster desktop sometimes slowed an agent: GPT-6 Astra’s task time rose from 89.5 to 99.0 seconds with Fast I/O enabled, averaged over five runs.
Combined views
91
1 Source, first seen ago