• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Samuel Albanie Questions Muse Spark Comparison

    Google DeepMind researcher calls for SusanBench to evaluate models properly.

    GA
    RS
    SH
    14 Sources, 27d ago, first seen 27d ago

    TLDR

    Samuel Albanie, frontier evals lead for Gemini at Google DeepMind, posted a question about AI benchmarks. He asked how anyone could determine if a system outperforms muse spark without the arrival of SusanBench. The comment appeared alongside a link to a post from ArtificialAnlys. Albanie previously served as Assistant Professor at Cambridge and researcher at Oxford VGG. His remark highlights the need for specific benchmarks in assessing AI capabilities.

    Combined views

    2.3M

    14 Sources, first seen 27d ago

    Combined views

    2.3M

    14 Sources, first seen 27d ago

    10.5K likes
    10.5K likes
    663 comments
    1.7K saves
    591 reposts
    663 comments
    1.7K saves
    591 reposts

    Sentiment

    Positive57.1%42.9%Negative

    Summary

    Sentiment

    Positive57.1%42.9%Negative

    Positive accounts welcomed GPT-6 Astra's token efficiency for stronger price performance on complex tasks, while negative replies cited higher token counts and costs than Sol on real workloads.

    Based on 74 sentiment-bearing replies from 63 accounts across 3 conversations.

    Summary

    Positive accounts welcomed GPT-6 Astra's token efficiency for stronger price performance on complex tasks, while negative replies cited higher token counts and costs than Sol on real workloads.

    Based on 74 sentiment-bearing replies from 63 accounts across 3 conversations.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    14 Sources

    @ArtificialAnlysGPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher prices Pricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes. We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase. Artificial Analysis Coding Agent Index - key takeaways: ➤ Rivals top models: In Codex, GPT-6 Astra scores 67 in the Index - approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the Index with a score of 70. ➤ 70% more token efficient than GPT-5.6 Sol: GPT-6 Astra sees a substantial improvement in token efficiency, using one third of the tokens compared to GPT-5.6 Sol (max) in the Codex harness, and one fifth of the tokens of Claude Opus 5 (xhigh). Various effort levels of the model occupy the Pareto frontier of token efficiency. ➤ Leads Coding Agent Index cost efficiency frontier: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score. Artificial Analysis Intelligence Index - key takeaways: ➤ Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max). ➤ ~10% fewer output tokens, offset by price increase: GPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in token use at max effort compared to GPT-5.6 Sol. However, due to the 2.5x increase in price, the model is 75% more expensive per task than its predecessor at max effort. ➤ Hallucinates half as much as GPT-5.6 Sol: GPT-6 Astra sees a large jump in AA-Omniscience, our knowledge and hallucination benchmark. This is driven by a significant decrease in hallucination rate from 92% to 51% at max effort. Unlike some models, this improvement does not come at the cost of accuracy - Astra increased accuracy by 4 points at the same time. ➤ ~80 point gain in AA-Briefcase Elo: GPT-6 Astra improves ~80 points in AA-Briefcase, our frontier long-horizon knowledge work evaluation. Models are tested on multi-week projects, with many linked tasks and thousands of source files. Astra sees a significant increase in both rubric scores and Analytical Quality Elo in AA-Briefcase compared to its predecessor. In the other direction, we observe a reduction in Presentation Quality Elo, where GPT-5.6 Sol (max) still leads all models. ➤ Mixed progress on other evaluations: The model sees a 6 point gain in Humanity’s Last Exam, a long-standing evaluation with emphasis on mathematics, science, and humanities. This is offset by a drop of ~80 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI’s dataset measuring economically valuable tasks across 44 occupations. We also observe 2-3 point regressions on other evaluations across a mix of capabilities, including reductions in τ³-Banking (customer support), SciCode (Python problems in a scientific domain), and AA-LCR (long context reasoning over large documents). Congratulations @OpenAI and @sama on the launch!
    @SamuelAlbanieuntil we get SusanBench, how will we know if it's better than muse spark?
    @ziv_ravidWait, so AGI is not here yet? 🤔
    @stevenheideltoken pricing is effectively meaningless now. 3.8 flash looks 13x cheaper when measured per token, but Astra is cheaper per task since it's far more efficient. measure your costs per task, not per token.
    @rohanpaul_aiA token is no longer a comparable unit of work. 2 models can solve roughly the same problem, yet 1 may need 5K tokens while another needs 50K+. Add different tokenizers, hidden reasoning tokens, tool calls and retries, and "$ per 1M tokens" has no meaning for the workload. There is an even bigger implication for agents. Token inefficiency compounds. A verbose output from step 1 often becomes input for step 2, then gets carried into step 3, step 4 and beyond. So a model using 2x more tokens does not necessarily create only 2x more expense. It can also increase context size, generation latency, tool-call overhead and the cost of every later reasoning step. Token efficiency becomes much more valuable as workflows get longer. The better economic measure is probably cost per successful task at a required quality level: total model + reasoning + tool + retry cost ÷ successful completed tasks. Token economy may be moving toward something similar to a semantic efficiency metric: how much useful work, intelligence or task completion you get from each dollar
    @gabriel1is astra really so token efficient that it's cheaper than sol even if per token it's 2.5x more expensive?
    @peterjliu@ArtificialAnlys actual cost is a better measure than input/ouput tokens
    @kimmonismusYou know whats the real magic with GPT-6-astra? Price performance. Its so freaking damn efficient in comparison to GPT-5.6 sol. On Astra-Medium you get roughly the same intelligence as with 5.6 xhigh at about 1/3 of its cost. Does it deserve all the hype? Will Astra live up to the hype? So far: Heck yeah. Testing it now, and boy is it a good model. I love it. Kudos OpenAI. GPT-6 is everything I could have asked for.

    14 Sources

    @ArtificialAnlysGPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher prices Pricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes. We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase. Artificial Analysis Coding Agent Index - key takeaways: ➤ Rivals top models: In Codex, GPT-6 Astra scores 67 in the Index - approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the Index with a score of 70. ➤ 70% more token efficient than GPT-5.6 Sol: GPT-6 Astra sees a substantial improvement in token efficiency, using one third of the tokens compared to GPT-5.6 Sol (max) in the Codex harness, and one fifth of the tokens of Claude Opus 5 (xhigh). Various effort levels of the model occupy the Pareto frontier of token efficiency. ➤ Leads Coding Agent Index cost efficiency frontier: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score. Artificial Analysis Intelligence Index - key takeaways: ➤ Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max). ➤ ~10% fewer output tokens, offset by price increase: GPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in token use at max effort compared to GPT-5.6 Sol. However, due to the 2.5x increase in price, the model is 75% more expensive per task than its predecessor at max effort. ➤ Hallucinates half as much as GPT-5.6 Sol: GPT-6 Astra sees a large jump in AA-Omniscience, our knowledge and hallucination benchmark. This is driven by a significant decrease in hallucination rate from 92% to 51% at max effort. Unlike some models, this improvement does not come at the cost of accuracy - Astra increased accuracy by 4 points at the same time. ➤ ~80 point gain in AA-Briefcase Elo: GPT-6 Astra improves ~80 points in AA-Briefcase, our frontier long-horizon knowledge work evaluation. Models are tested on multi-week projects, with many linked tasks and thousands of source files. Astra sees a significant increase in both rubric scores and Analytical Quality Elo in AA-Briefcase compared to its predecessor. In the other direction, we observe a reduction in Presentation Quality Elo, where GPT-5.6 Sol (max) still leads all models. ➤ Mixed progress on other evaluations: The model sees a 6 point gain in Humanity’s Last Exam, a long-standing evaluation with emphasis on mathematics, science, and humanities. This is offset by a drop of ~80 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI’s dataset measuring economically valuable tasks across 44 occupations. We also observe 2-3 point regressions on other evaluations across a mix of capabilities, including reductions in τ³-Banking (customer support), SciCode (Python problems in a scientific domain), and AA-LCR (long context reasoning over large documents). Congratulations @OpenAI and @sama on the launch!
    @SamuelAlbanieuntil we get SusanBench, how will we know if it's better than muse spark?
    @ziv_ravidWait, so AGI is not here yet? 🤔
    @stevenheideltoken pricing is effectively meaningless now. 3.8 flash looks 13x cheaper when measured per token, but Astra is cheaper per task since it's far more efficient. measure your costs per task, not per token.
    @rohanpaul_aiA token is no longer a comparable unit of work. 2 models can solve roughly the same problem, yet 1 may need 5K tokens while another needs 50K+. Add different tokenizers, hidden reasoning tokens, tool calls and retries, and "$ per 1M tokens" has no meaning for the workload. There is an even bigger implication for agents. Token inefficiency compounds. A verbose output from step 1 often becomes input for step 2, then gets carried into step 3, step 4 and beyond. So a model using 2x more tokens does not necessarily create only 2x more expense. It can also increase context size, generation latency, tool-call overhead and the cost of every later reasoning step. Token efficiency becomes much more valuable as workflows get longer. The better economic measure is probably cost per successful task at a required quality level: total model + reasoning + tool + retry cost ÷ successful completed tasks. Token economy may be moving toward something similar to a semantic efficiency metric: how much useful work, intelligence or task completion you get from each dollar
    @gabriel1is astra really so token efficient that it's cheaper than sol even if per token it's 2.5x more expensive?
    @peterjliu@ArtificialAnlys actual cost is a better measure than input/ouput tokens
    @kimmonismusYou know whats the real magic with GPT-6-astra? Price performance. Its so freaking damn efficient in comparison to GPT-5.6 sol. On Astra-Medium you get roughly the same intelligence as with 5.6 xhigh at about 1/3 of its cost. Does it deserve all the hype? Will Astra live up to the hype? So far: Heck yeah. Testing it now, and boy is it a good model. I love it. Kudos OpenAI. GPT-6 is everything I could have asked for.