• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    GPT-6 Astra scored 14% on MazeBench, a user reports

    A user says Astra spent 60-plus hours on the 3D open-world spatial reasoning evaluation.

    SB
    YA
    AC
    6 Sources, 23d ago, first seen 23d ago

    TLDR

    A user reports that GPT-6 Astra spent 60-plus hours on MazeBench and finished with a 14% score. They describe MazeBench as a 3D open-world evaluation of spatial reasoning.

    Combined views

    464.3K

    6 Sources, first seen 23d ago

    3.9K likes

    Combined views

    464.3K

    6 Sources, first seen 23d ago

    3.9K likes
    223 comments
    765 saves
    273 reposts
    223 comments
    765 saves
    273 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    6 Sources

    @patience_caveMazeBench vs GPT-6 Astra Astra spent 60+ hours in this 3D open world spatial reasoning eval. Final score: 14%
    @AndrewCurran_An update from a benchmark I trust more than 99% of existing ones shows a massive jump with Astra. As capability advancement grows ever more jagged, users increasingly see only the parts of the elephant that relate to their tasks.
    @kimmonismusAnother exciting reasoning benchmark, MazeBench, where GPT-Astra outperforms any other model, especially Fable 5.1. At this point it’s clear to me that Astra is by far the best model. And since we already know that OpenAI is close to release another model, I couldn’t be more exited.
    @SebastienBubeckRT @AndrewCurran_: An update from a benchmark I trust more than 99% of existing ones shows a massive jump with Astra. As capability advance…
    @rohanpaul_aiBig score by GPT-6 Astra, 14% on MazeBench without Python, 7X Claude Fable 5.1's 2%. MazeBench specifically stresses long-horizon spatial reasoning, with puzzles that can require more than 100 moves and repeated interaction with a changing environment. So Astra seems able to hold a long plan in its head and execute it without relying on Python to solve the environment for it Its remaining failures on genuinely 3D puzzles show that the improvement is mainly in sustained planning, not complete spatial understanding. MazeBench hides 100 gems across more than 200 rooms, forcing agents to rotate views, manipulate boxes, recover from mistakes, and execute plans that can exceed 100 moves. The no-Python track makes it so much harder because code-enabled agents can reverse-engineer the physics and run solver algorithms instead of carrying the full spatial plan themselves. MazeBench's creator says Astra still surpassed GPT-5.6 Sol's 13% code-enabled run and often planned 10-20 moves ahead in batched actions. That batching reportedly reduced a projected 3B-token trajectory to about 350M tokens, but the run still lasted more than 60 hours.
    @yoavartziAstra does seem like a step change in spatial reasoning. Impressive....

    6 Sources

    @patience_caveMazeBench vs GPT-6 Astra Astra spent 60+ hours in this 3D open world spatial reasoning eval. Final score: 14%
    @AndrewCurran_An update from a benchmark I trust more than 99% of existing ones shows a massive jump with Astra. As capability advancement grows ever more jagged, users increasingly see only the parts of the elephant that relate to their tasks.
    @kimmonismusAnother exciting reasoning benchmark, MazeBench, where GPT-Astra outperforms any other model, especially Fable 5.1. At this point it’s clear to me that Astra is by far the best model. And since we already know that OpenAI is close to release another model, I couldn’t be more exited.
    @SebastienBubeckRT @AndrewCurran_: An update from a benchmark I trust more than 99% of existing ones shows a massive jump with Astra. As capability advance…
    @rohanpaul_aiBig score by GPT-6 Astra, 14% on MazeBench without Python, 7X Claude Fable 5.1's 2%. MazeBench specifically stresses long-horizon spatial reasoning, with puzzles that can require more than 100 moves and repeated interaction with a changing environment. So Astra seems able to hold a long plan in its head and execute it without relying on Python to solve the environment for it Its remaining failures on genuinely 3D puzzles show that the improvement is mainly in sustained planning, not complete spatial understanding. MazeBench hides 100 gems across more than 200 rooms, forcing agents to rotate views, manipulate boxes, recover from mistakes, and execute plans that can exceed 100 moves. The no-Python track makes it so much harder because code-enabled agents can reverse-engineer the physics and run solver algorithms instead of carrying the full spatial plan themselves. MazeBench's creator says Astra still surpassed GPT-5.6 Sol's 13% code-enabled run and often planned 10-20 moves ahead in batched actions. That batching reportedly reduced a projected 3B-token trajectory to about 350M tokens, but the run still lasted more than 60 hours.
    @yoavartziAstra does seem like a step change in spatial reasoning. Impressive....