Announcement
cua-speedrun benchmarks computer-use agents on accuracy, speed and cost
Its creators say it standardizes hardware, desktops, tasks and timing rules to compare agents across four benchmarks.
TLDR
The cua-speedrun team says its benchmark measures accuracy, task time and cost under consistent conditions. On OSWorld, it reports that Kimi K3 at max effort and GPT-6 Astra at high effort both scored 89.6%, but Kimi took 4.4 times as long per task. The team also says more reasoning effort shortened task time for some models.
Combined views
5.6K
6 Sources, first seen 2h ago
