• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Sai's claimed 79.38% partial-success rate ranks first on OSWorld

    Sai's team says OSWorld verified its run and cites a $14.34 cost per task, versus $24.11 for Anthropic's Claude Computer Use with Opus 5.

    XE
    AL
    2 Sources, 2h ago, first seen ago

    TLDR

    In an October 5 post, Sai's team said OSWorld verified its run and that Sai ranked first in partial success rate at 79.38% across real-world desktop environments. It reported a $14.34 cost per task, versus $24.11 for Anthropic's Claude Computer Use with Opus 5, and claimed higher overall task completion.

    Combined views

    426

    2 Sources, first seen 2h ago

    Combined views

    426

    2 Sources, first seen 2h ago

    10 likes
    2h
    10 likes
    2 comments
    4 saves
    3 reposts
    2 comments
    4 saves
    3 reposts
    Featured Source

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @angli_aiToday, OSWorld officially verified our results on their leaderboard: Sai @sai_borg is #1 in partial success rate (79.38%) across real-world desktop environments. What I'm most proud of isn't just the accuracy, but the efficiency: Cost per task: $14.34 vs. Anthropic's Claude Computer Use with Opus 5 at $24.11 That’s higher overall task completion while slashing the per-task inference cost by over 40%. Getting frontier models to reliably drive an operating system isn't just about throwing raw compute at the screen — it's about the orchestration layer, grounding, and interaction primitives underneath. Huge thanks to the OSWorld team for setting the gold standard benchmark and verifying our run.2h
    @xwang_lkRT @angli_ai: Today, OSWorld officially verified our results on their leaderboard: Sai @sai_borg is #1 in partial success rate (79.38%) ac…2h

    2 Sources

    @angli_aiToday, OSWorld officially verified our results on their leaderboard: Sai @sai_borg is #1 in partial success rate (79.38%) across real-world desktop environments. What I'm most proud of isn't just the accuracy, but the efficiency: Cost per task: $14.34 vs. Anthropic's Claude Computer Use with Opus 5 at $24.11 That’s higher overall task completion while slashing the per-task inference cost by over 40%. Getting frontier models to reliably drive an operating system isn't just about throwing raw compute at the screen — it's about the orchestration layer, grounding, and interaction primitives underneath. Huge thanks to the OSWorld team for setting the gold standard benchmark and verifying our run.2h
    @xwang_lkRT @angli_ai: Today, OSWorld officially verified our results on their leaderboard: Sai @sai_borg is #1 in partial success rate (79.38%) ac…2h