Sai Agent Hits 73% on OSWorld 2.0
Simular's Sai agent outperformed GPT models on workplace tasks at lower cost.
TLDR
The New Stack reported that Simular's Sai agent reached 73% on the OSWorld 2.0 benchmark for routine workplace tasks. The coverage states Sai topped GPT-5.6 Sol and Opus 5 while running at roughly two-thirds the cost. A headline quote in the post reads “Posterity will find it ludicrous.” The article frames the result as performance on real but necessary work rather than novel challenges. No further confirmation or independent test results appear in the supplied lines.
Combined views
691
1 Source, first seen 26d ago