• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI

    Psyho Reports Astra Solved Every Logic Puzzle Tested

    Top competitive programmer tested Astra on hard unpublished logic puzzles.

    BP
    PS
    2 Sources, 24d ago, first seen 24d ago

    TLDR

    Przemysław Dębiak, known as Psyho and a former OpenAI engineer, posted that he gave Astra many logic puzzles to solve without writing code. He stated the system completed every one, covering puzzles too difficult for world championships, very large grids, unpublished examples, and those impossible to solve by backtracking alone. The post compared performance to an earlier solver version. The statement reflects Dębiak's account of his own tests, as shared in the visible tweet.

    Combined views

    98.1K

    2 Sources, first seen 24d ago

    Combined views

    98.1K

    2 Sources, first seen 24d ago

    2K likes
    2K likes
    52 comments
    347 saves
    170 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    52 comments
    347 saves
    170 reposts

    Sentiment

    Positive——Negative

    Summary

    Not enough discussion yet.

    No sentiment analysis available yet.

    Today's Rank

    —

    Not ranked yet

    Today's Rank

    —

    Not ranked yet

    2 Sources

    @FakePsyhoI threw a ton of various logic puzzles at Astra and... it logically (no code) solved ALL of them. I mean literally every puzzle that I tried, including: too hard for world championship, very large grids, unpublished puzzles, puzzles that are impossible to solve via backtracking alone. For comparison, sol 5.6 was around 20-30%. I analyzed the results and all of these were legit. Honestly I don't care if OpenAI RLed logic puzzles to death. Astra is now better at explaining solving paths than I am.
    @BorisMPowerWould love for this to become an eval, especially if there’s a way to create harder puzzles so it’s not saturated by Astra

    2 Sources

    @FakePsyhoI threw a ton of various logic puzzles at Astra and... it logically (no code) solved ALL of them. I mean literally every puzzle that I tried, including: too hard for world championship, very large grids, unpublished puzzles, puzzles that are impossible to solve via backtracking alone. For comparison, sol 5.6 was around 20-30%. I analyzed the results and all of these were legit. Honestly I don't care if OpenAI RLed logic puzzles to death. Astra is now better at explaining solving paths than I am.
    @BorisMPowerWould love for this to become an eval, especially if there’s a way to create harder puzzles so it’s not saturated by Astra