• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Muse Spark 1.3 reportedly lifts average test pass rate from 47.0% to 68.6% at xhigh

    A user comparing versions 1.1 and 1.3, both at xhigh, says 1.3 used 39% fewer tokens across the benchmark and brought the number of programs at least 95% resolved from eight to 33.

    ZS
    JY
    JM
    12 Sources, ,

    TLDR

    A user reports that Muse Spark’s average test pass rate rose from 47.0% in version 1.1 to 68.6% in 1.3 at xhigh, while token use fell 39% across the benchmark. In a separate comparison, they say Muse Spark at max averaged $6.46 per task and fully resolved five programs, versus GPT-5.6 Sol xhigh at $6.08 and two. They put Opus 5 at nine fully resolved programs and $50.53 per task.

    Combined views

    10.2K

    12 Sources, first seen 14h ago

    Combined views

    10.2K

    12 Sources, first seen 14h ago

    128 likes
    14h ago
    first seen 14h ago
    128 likes
    9 comments
    10 saves
    35 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Featured Source
    9 comments
    10 saves
    35 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    12 Sources

    @jyangballinSo excited to finally share: Meta Muse Spark 1.[1/2/3] on ProgramBench In PB, a SWE-agent must write a whole program (sqlite, ffmpeg, php) from scratch Muse Spark 1.3 takes #2 (max) and #3 (xhigh) overall at incredible cost tiers Intelligence too cheap to meter at its finest
    @18jeffreymacan’t believe this lil guy is #2 on our leaderboard check out @jyangballin’s thread for some awesome analysis!
    @parth007_96Checkout our latest update on programbench leaderboard Muse is really good at difficult long horizon tasks
    @KLieretProgramBench leaderboard updates: Muse Spark 1.3 significantly outperforming GPT 5.6 Sol. Opus 5 is 1st place, but also costs 8x more
    @magpie_rayhouMuseSpark improves really rapidly in the past months on coding and long horizon tasks — ProgramBench shows it here. More to come.
    @EdwardSun0909RT @jyangballin: So excited to finally share: Meta Muse Spark 1.[1/2/3] on ProgramBench In PB, a SWE-agent must write a whole program (sql…

    12 Sources

    @jyangballinSo excited to finally share: Meta Muse Spark 1.[1/2/3] on ProgramBench In PB, a SWE-agent must write a whole program (sqlite, ffmpeg, php) from scratch Muse Spark 1.3 takes #2 (max) and #3 (xhigh) overall at incredible cost tiers Intelligence too cheap to meter at its finest
    @18jeffreymacan’t believe this lil guy is #2 on our leaderboard check out @jyangballin’s thread for some awesome analysis!
    @parth007_96Checkout our latest update on programbench leaderboard Muse is really good at difficult long horizon tasks
    @KLieretProgramBench leaderboard updates: Muse Spark 1.3 significantly outperforming GPT 5.6 Sol. Opus 5 is 1st place, but also costs 8x more
    @magpie_rayhouMuseSpark improves really rapidly in the past months on coding and long horizon tasks — ProgramBench shows it here. More to come.
    @EdwardSun0909RT @jyangballin: So excited to finally share: Meta Muse Spark 1.[1/2/3] on ProgramBench In PB, a SWE-agent must write a whole program (sql…