Sai Agent Hits 73% on OSWorld 2.0
Simular's Sai agent outperformed GPT models on workplace tasks at lower cost.
Simular's Sai agent outperformed GPT models on workplace tasks at lower cost.
The New Stack reported that Simular's Sai agent reached 73% on the OSWorld 2.0 benchmark for routine workplace tasks. The coverage states Sai topped GPT-5.6 Sol and Opus 5 while running at roughly two-thirds the cost. A headline quote in the post reads “Posterity will find it ludicrous.” The article frames the result as performance on real but necessary work rather than novel challenges. No further confirmation or independent test results appear in the supplied lines.
667
1 post, first seen 21h ago
Simular's Sai agent outperformed GPT models on workplace tasks at lower cost.
The New Stack reported that Simular's Sai agent reached 73% on the OSWorld 2.0 benchmark for routine workplace tasks. The coverage states Sai topped GPT-5.6 Sol and Opus 5 while running at roughly two-thirds the cost. A headline quote in the post reads “Posterity will find it ludicrous.” The article frames the result as performance on real but necessary work rather than novel challenges. No further confirmation or independent test results appear in the supplied lines.