• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Theo Says Astra Advances Make Benchmarks Harder

    Creator Theo Browne reacts to capabilities shown by the Astra system.

    T-
    1 Source, 27d ago, first seen 27d ago

    TLDR

    Theo Browne, the software developer and YouTuber behind the T3 Stack, posted a reply about the Astra system. He stated that many areas where Astra performs better involve tasks that were not worth benchmarking before. Browne described the change as a huge leap in things he did not expect large language models to achieve. He added that benchmarks will become much harder to create going forward. The comment appears in a reply visible among posts discussing the system and its results.

    Combined views

    50.2K

    1 Source, first seen 27d ago

    Combined views

    50.2K

    1 Source, first seen 27d ago

    939 likes
    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    939 likes
    26 comments
    40 saves
    9 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    26 comments
    40 saves
    9 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    1 Source

    @theoTo be fair, a significant # of the things Astra is better at are things that weren’t even worth benchmarking before. It’s a huge leap in shit I didn’t think LLMs could ever do. Benchmarks are going to be way harder to make going forward

    1 Source

    @theoTo be fair, a significant # of the things Astra is better at are things that weren’t even worth benchmarking before. It’s a huge leap in shit I didn’t think LLMs could ever do. Benchmarks are going to be way harder to make going forward