• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Leo Linsky Claims GPT-6-Astra Leads Multi-Agent Tests

    Tweet shares results from testing in complex multi-agent coding environments.

    LL
    1 Source, 25d ago, first seen 25d ago

    TLDR

    Leo Linsky posted that his evaluation of GPT-6-Astra covered 100 complex multi-agent coding environments with competing and cooperating models in open-ended tasks. He stated the model is the new frontier leader by a wide margin and more dominant than the Fable 5 release. The post notes it outperforms the second-best model and includes an attachment showing the GBENCH Intelligence Benchmark leaderboard table.

    Combined views

    105.3K

    1 Source, first seen 25d ago

    Combined views

    105.3K

    1 Source, first seen 25d ago

    1.2K likes
    1.2K likes
    43 comments
    339 saves
    103 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    43 comments
    339 saves
    103 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    1 Source

    @leo_linskyWe evaluated GPT-6-Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks. It's the new frontier model by a landslide. It's even more dominant than the Fable 5 release, because not only does it wipe the floor with the second best model (Fable 5.1), it was also 80% cheaper and 30% faster in agentic coding. This is a groundbreaking model. The biggest breakthrough since Opus 4.5, maybe even since GPT 4. This is also a testament to how unsaturated our evals are. Some of the open-ended environments we run were introduced almost a year ago and can still compete frontier model submissions against older, deprecated models like Sonnet 4 without signs of saturation. Our newer ones are much more challenging and completely auto-generated. We are clearly living in a post-AGI world. Do something interesting and useful with it.

    1 Source

    @leo_linskyWe evaluated GPT-6-Astra in 100 complex, unsaturated multi-agent coding environments, competing and cooperating with other models in open-ended tasks. It's the new frontier model by a landslide. It's even more dominant than the Fable 5 release, because not only does it wipe the floor with the second best model (Fable 5.1), it was also 80% cheaper and 30% faster in agentic coding. This is a groundbreaking model. The biggest breakthrough since Opus 4.5, maybe even since GPT 4. This is also a testament to how unsaturated our evals are. Some of the open-ended environments we run were introduced almost a year ago and can still compete frontier model submissions against older, deprecated models like Sonnet 4 without signs of saturation. Our newer ones are much more challenging and completely auto-generated. We are clearly living in a post-AGI world. Do something interesting and useful with it.