• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    GPT-6 Astra Sets ARC-AGI-3 Record

    François Chollet and ARC Prize discuss GPT-6 Astra's benchmark results and what they mean for AGI claims.

    GB
    FC
    ('
    68 Sources, 27d ago, first seen 27d ago

    TLDR

    ARC Prize posted that OpenAI's GPT-6 Astra scores 63 percent on ARC-AGI-3 and reaches 99 percent with a new harness. The system surpasses human performance on 96 percent of tasks and builds precise symbolic models of novel environments. Matt Mazur reported 62.7 percent using their standard harness, more than doubling the prior verified score. François Chollet stated that benchmarking co-evolves with models and that saturating ARC-AGI-3 does not prove AGI. Mike Knoop called the result a large leap toward AGI but noted open-ended invention remains unsolved. Chollet added that progress arrived twice as fast as his earlier one-year estimate.

    Combined views

    4.6M

    68 Sources, first seen 27d ago

    Combined views

    4.6M

    68 Sources, first seen 27d ago

    41.3K likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    41.3K likes
    2.1K comments
    6K saves
    4.3K reposts
    2.1K comments
    6K saves
    4.3K reposts

    Sentiment

    Positive48.9%51.1%Negative

    Summary

    Sentiment

    Positive48.9%51.1%Negative

    Positive accounts praised GPT-6 Astra’s ARC-AGI-3 results as amazing progress with strong efficiency and no performance wall, while negative replies called the benchmark meaningless or low-utility after high scores.

    Based on 213 sentiment-bearing replies from 188 accounts across 10 conversations.

    Summary

    Positive accounts praised GPT-6 Astra’s ARC-AGI-3 results as amazing progress with strong efficiency and no performance wall, while negative replies called the benchmark meaningless or low-utility after high scores.

    Based on 213 sentiment-bearing replies from 188 accounts across 10 conversations.

    68 Sources

    @arcprizeGPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis:
    @fcholletGPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation. Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence. Read our post on Astra and what these results mean: https://arcprize.org/blog/astra
    @teortaxesTexwill there be ARC-AGI-4? Or are we done?
    @willdepueit's over boys. i think we're out of ARCs to saturate
    @sanderstedRT @arcprize: GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - I…
    @GregKamradtReflections on Astra from a benchmark perspective: 1. No harness was used in the making of these scores Up until this point we’ve seen multiple groups report results above 90% on ARC-AGI-3. Examples include PRO-LONG, Tycho, Schema, Prime Agent, see community leaderboard [1] for the rest. They all share a general shape: Lossless Memory + Programmatic Analysis + Explicit Hypothesis Testing + Persistent State + Cheap Internal Compute. We’ve predicted (see v3 paper [2]) that as models get better they would _subsume_ more of this harness functionality. Astra is another step in that direction. Its achieves high scores without a heavy client-side harness. It externalizes its working model in its outputs as compact natural-language programs. If this trend continues (I anticipate it will), more of that structure will become latent computation inside the model. We will see fewer intermediate steps and be limited to analyzing the resulting behavior [3]. From the developers perspective, our lack of visibility has already started with opaque/encrypted reasoning logs. How did the model reason through the game? We aren’t exactly sure other than what it chose to put in its output. 2. ARC-AGI-3 Astra's Human efficiency Prior to launching v3, we hypothesized that counting the # of actions that AI took to solve a game would show a dividing line between humans and AI. While this is true of brute-force approaches, we’ve found that frontier LLMs either “get it” or they don’t. If they _do_ get the mechanics of a v3 game, then they’ll complete it within range of human action efficiency. Astra uses fewer actions than humans on 96% of levels. In alternative world it might have used 2-3x the actions for the same result. 3. Close-ended environments ARC-AGI-3 was designed as series of closed-ended environments. They are short, deterministic games that humans can do in 5-10 minutes. To complete v3, you must translate the environment into a world model and then anticipate ahead. The envs can be reduced to simple programs. However, reality is not a simple program, but humans are still able to abstract reality's complexity into useful value. Future benchmarks should allow for much more open endedness and innovation. 4. Benchmark tension There is a healthy tension for us as a benchmarking organization between supporting researchers and testing frontier models. Using the same benchmark for both aims means trade offs. For example, how much of the benchmark do we make public? Any signal we give about a benchmark’s domain invites developer aware targeting (whether implicit or explicit). If you have thoughts about this, please reach out. 5. Measuring General Intelligence How can you assert humans are generally intelligent? The 2x ways I feel most confident in is 1) Our collective technological progress over the last thousands of years and 2) The fact that a newborn can grow up and participate in the world of 1K years ago and the world of today. Both of those are long-horizon (years/millennia) tests. Trying to find a “unit test” of generalization that we can test a model with in a few days is the work! 6. This is a step towards generalization, not sufficient evidence of AGI.
    @scaling01ARC-AGI-3 - "GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels"
    @mikeknoopGPT-6 Astra is the new SOTA on ARC-AGI-3 It's a qualitatively large leap towards AGI and the pace of progress is frankly surprising. That said, we lack evidence to call this AGI yet. While we are still studying the human capability gaps, we believe open-ended invention is unsolved, and this will form the new basis for ARC-AGI-4. Reminder that ARC v3 tests whether models can figure out how to efficiently make sense of unfamiliar environments and operate autonomously toward goals within them. Summary: properly equipped Astra can now do this. Astra direct model scored 66% at ~$500/game. Up from Sol 8%. This is apples-to-apples with all other ARC v3 scores we've verified to date. This score gives the best look at the general intelligence of base Astra. Using an OpenAI-specific adapter (which is open source) we verified the first ~100% at $300/game. If you're using Astra via API I would prioritize migrating to /conversations endpoint with compaction enabled. Direct model testing continues to carry important scientific interest towards AGI. But provider adapters are a better estimate of what you should expect day-to-day when using these models inside of products. The action efficiency of Astra is also pretty incredible. With the provider adapter, Astra used 50% fewer actions on average than our human baseline (who both score 100%). A surprise for us is that across all models, action efficiency on v3 is very bi-modal. Models either get the game or not. Weaker agent models aren't able to "brute force" their way to an inefficient understanding. Tactically, our results support always using Astra's higher reasoning tiers when deploying agents into environments. It not only makes Astra better but is cheaper too (as a result of needing less turns). The key technique Astra seems to leverage is on-the-fly symbolic world modeling. It develops a custom algebra and notation for the environment through interaction and uses this understanding to plan (all in token space). This approach is similar to earlier harnesses we saw which created symbolic models through side-car executable programs to model the environment and plan. Based on Astra results I now see clear paths to recursive self improvement via horizontal data scaling and continual learning where the learned state is information stored and managed outside of the model weights (text, programs, databases, etc). I expect more and more "learning" to happen outside the model weights and this will be of growing interest to both researcher and AI product builders. One remaining area of important capability we haven't seen demonstrated yet by any model is open-ended invention and discovery. Humans clearly posses this capability and we've so far seen zero evidence that frontier AI -- even as powerful as Astra -- is capable of this feat. Consider AI research 10 years ago vs today. Humans invented so much: transformer, self-supervised LLM pretraining, diffusion, RLHF, test time adaptation, ... Society needs AI systems that not only can autonomously operate towards defined goals. We need AI systems help craft great futures that are inherently undefined. I believe we in at a very tenuous spot. We now have extremely powerful automation system that capable of causing a lot of disruption. Without the counterbalancing force of demonstrated invention by AI, I fear society will turn more and more against open frontier AI research and development -- trapping us in a very bad middle ground. We must move as swiftly and with focus towards this important AI capability. And this has become the primary focus of our work on ARC-AGI.
    @thsottiauxWe are going to need a different AGI benchmark. Where is the goalpost moving next?
    @mhmazurAnalyzing GPT-6 Astra's performance on ARC-AGI has been one of the highlights of my career: With our standard harness - which lets models carry notes forward and is what we've historically used for apples-to-apples comparisons - Astra more than doubles the previous verified high score, reaching 62.7%. We also tested Astra with a new provider adapter harness, configured to preserve the model's opaque reasoning state between requests and use native compaction as the context window grows. With this setup, Astra scores a whopping 99.9% on ARC-AGI-3. Going forward, we'll test models with both the standard harness and a provider adapter harness whenever possible. On ARC-AGI-2, Astra sets a new state-of-the-art score of 95.0%. On ARC-AGI-1, it ties Fable 5 at 98.5%. In the standard harness, we also saw the cost/performance curve bend backward for the first time: the max run cost $26K, versus $38K for low and $48K for medium, because it completed games in fewer actions. In the max provider adapter run, Astra used fewer actions than the median tested human on 96% of completed levels. Its replays showed a combination of: - Persistent world models: it tracked objects, positions, orientations, and rules across turns - Coordinate abstraction: it converted pixels into logical cells and reused coordinate formulas - Long-horizon planning: it tracked routes, subgoals, constraints, and expected state changes - Cumulative learning: it retained discoveries and discarded failed hypotheses - Checkpointed recovery: it used resets to revise plans without repeating exploration In short, it represents a huge leap in capabilities over what came before it. If you have any questions about Astra's performance on any of the ARC-AGI benchmarks, let me know. I'm happy to dig in and share any details I can.

    68 Sources

    @arcprizeGPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis:
    @fcholletGPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation. Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself. We see Astra as a major breakthrough in model intelligence. Read our post on Astra and what these results mean: https://arcprize.org/blog/astra
    @teortaxesTexwill there be ARC-AGI-4? Or are we done?
    @willdepueit's over boys. i think we're out of ARCs to saturate
    @sanderstedRT @arcprize: GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - I…
    @GregKamradtReflections on Astra from a benchmark perspective: 1. No harness was used in the making of these scores Up until this point we’ve seen multiple groups report results above 90% on ARC-AGI-3. Examples include PRO-LONG, Tycho, Schema, Prime Agent, see community leaderboard [1] for the rest. They all share a general shape: Lossless Memory + Programmatic Analysis + Explicit Hypothesis Testing + Persistent State + Cheap Internal Compute. We’ve predicted (see v3 paper [2]) that as models get better they would _subsume_ more of this harness functionality. Astra is another step in that direction. Its achieves high scores without a heavy client-side harness. It externalizes its working model in its outputs as compact natural-language programs. If this trend continues (I anticipate it will), more of that structure will become latent computation inside the model. We will see fewer intermediate steps and be limited to analyzing the resulting behavior [3]. From the developers perspective, our lack of visibility has already started with opaque/encrypted reasoning logs. How did the model reason through the game? We aren’t exactly sure other than what it chose to put in its output. 2. ARC-AGI-3 Astra's Human efficiency Prior to launching v3, we hypothesized that counting the # of actions that AI took to solve a game would show a dividing line between humans and AI. While this is true of brute-force approaches, we’ve found that frontier LLMs either “get it” or they don’t. If they _do_ get the mechanics of a v3 game, then they’ll complete it within range of human action efficiency. Astra uses fewer actions than humans on 96% of levels. In alternative world it might have used 2-3x the actions for the same result. 3. Close-ended environments ARC-AGI-3 was designed as series of closed-ended environments. They are short, deterministic games that humans can do in 5-10 minutes. To complete v3, you must translate the environment into a world model and then anticipate ahead. The envs can be reduced to simple programs. However, reality is not a simple program, but humans are still able to abstract reality's complexity into useful value. Future benchmarks should allow for much more open endedness and innovation. 4. Benchmark tension There is a healthy tension for us as a benchmarking organization between supporting researchers and testing frontier models. Using the same benchmark for both aims means trade offs. For example, how much of the benchmark do we make public? Any signal we give about a benchmark’s domain invites developer aware targeting (whether implicit or explicit). If you have thoughts about this, please reach out. 5. Measuring General Intelligence How can you assert humans are generally intelligent? The 2x ways I feel most confident in is 1) Our collective technological progress over the last thousands of years and 2) The fact that a newborn can grow up and participate in the world of 1K years ago and the world of today. Both of those are long-horizon (years/millennia) tests. Trying to find a “unit test” of generalization that we can test a model with in a few days is the work! 6. This is a step towards generalization, not sufficient evidence of AGI.
    @scaling01ARC-AGI-3 - "GPT-6 Astra surpasses the human baseline in action efficiency on ARC-AGI-3. It used fewer actions than the median tested human on 96% of levels"
    @mikeknoopGPT-6 Astra is the new SOTA on ARC-AGI-3 It's a qualitatively large leap towards AGI and the pace of progress is frankly surprising. That said, we lack evidence to call this AGI yet. While we are still studying the human capability gaps, we believe open-ended invention is unsolved, and this will form the new basis for ARC-AGI-4. Reminder that ARC v3 tests whether models can figure out how to efficiently make sense of unfamiliar environments and operate autonomously toward goals within them. Summary: properly equipped Astra can now do this. Astra direct model scored 66% at ~$500/game. Up from Sol 8%. This is apples-to-apples with all other ARC v3 scores we've verified to date. This score gives the best look at the general intelligence of base Astra. Using an OpenAI-specific adapter (which is open source) we verified the first ~100% at $300/game. If you're using Astra via API I would prioritize migrating to /conversations endpoint with compaction enabled. Direct model testing continues to carry important scientific interest towards AGI. But provider adapters are a better estimate of what you should expect day-to-day when using these models inside of products. The action efficiency of Astra is also pretty incredible. With the provider adapter, Astra used 50% fewer actions on average than our human baseline (who both score 100%). A surprise for us is that across all models, action efficiency on v3 is very bi-modal. Models either get the game or not. Weaker agent models aren't able to "brute force" their way to an inefficient understanding. Tactically, our results support always using Astra's higher reasoning tiers when deploying agents into environments. It not only makes Astra better but is cheaper too (as a result of needing less turns). The key technique Astra seems to leverage is on-the-fly symbolic world modeling. It develops a custom algebra and notation for the environment through interaction and uses this understanding to plan (all in token space). This approach is similar to earlier harnesses we saw which created symbolic models through side-car executable programs to model the environment and plan. Based on Astra results I now see clear paths to recursive self improvement via horizontal data scaling and continual learning where the learned state is information stored and managed outside of the model weights (text, programs, databases, etc). I expect more and more "learning" to happen outside the model weights and this will be of growing interest to both researcher and AI product builders. One remaining area of important capability we haven't seen demonstrated yet by any model is open-ended invention and discovery. Humans clearly posses this capability and we've so far seen zero evidence that frontier AI -- even as powerful as Astra -- is capable of this feat. Consider AI research 10 years ago vs today. Humans invented so much: transformer, self-supervised LLM pretraining, diffusion, RLHF, test time adaptation, ... Society needs AI systems that not only can autonomously operate towards defined goals. We need AI systems help craft great futures that are inherently undefined. I believe we in at a very tenuous spot. We now have extremely powerful automation system that capable of causing a lot of disruption. Without the counterbalancing force of demonstrated invention by AI, I fear society will turn more and more against open frontier AI research and development -- trapping us in a very bad middle ground. We must move as swiftly and with focus towards this important AI capability. And this has become the primary focus of our work on ARC-AGI.
    @thsottiauxWe are going to need a different AGI benchmark. Where is the goalpost moving next?
    @mhmazurAnalyzing GPT-6 Astra's performance on ARC-AGI has been one of the highlights of my career: With our standard harness - which lets models carry notes forward and is what we've historically used for apples-to-apples comparisons - Astra more than doubles the previous verified high score, reaching 62.7%. We also tested Astra with a new provider adapter harness, configured to preserve the model's opaque reasoning state between requests and use native compaction as the context window grows. With this setup, Astra scores a whopping 99.9% on ARC-AGI-3. Going forward, we'll test models with both the standard harness and a provider adapter harness whenever possible. On ARC-AGI-2, Astra sets a new state-of-the-art score of 95.0%. On ARC-AGI-1, it ties Fable 5 at 98.5%. In the standard harness, we also saw the cost/performance curve bend backward for the first time: the max run cost $26K, versus $38K for low and $48K for medium, because it completed games in fewer actions. In the max provider adapter run, Astra used fewer actions than the median tested human on 96% of completed levels. Its replays showed a combination of: - Persistent world models: it tracked objects, positions, orientations, and rules across turns - Coordinate abstraction: it converted pixels into logical cells and reused coordinate formulas - Long-horizon planning: it tracked routes, subgoals, constraints, and expected state changes - Cumulative learning: it retained discoveries and discarded failed hypotheses - Checkpointed recovery: it used resets to revise plans without repeating exploration In short, it represents a huge leap in capabilities over what came before it. If you have any questions about Astra's performance on any of the ARC-AGI benchmarks, let me know. I'm happy to dig in and share any details I can.