• Home
  • Technology
  • Gaming
  • Entertainment
  • World & Business
  • Science
  • Sports
  • AI
HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
  • HomeTechnologyGamingEntertainmentWorld & BusinessScienceSportsAI
    • Home
    • Technology
    • Gaming
    • Entertainment
    • World & Business
    • Science
    • Sports
    • AI
    AI
    Announcement

    Ataraxos AI reportedly beats Stratego champion with 15 wins in 20 games

    A Nature paper reports 15 wins, one loss and four draws against Pim Niemeijer. The authors call it a superhuman result, but did not run a direct match against DeepNash.

    NB
    ZK
    EV
    10 Sources, ,

    TLDR

    A Nature paper reports that Ataraxos beat Stratego champion Pim Niemeijer with 15 wins, one loss and four draws over three weeks. The system combines self-play learning, a model of hidden pieces and search at decision time. Its 85% effective win rate counts draws as half-wins. The researchers did not compare it directly against DeepNash.

    Combined views

    252.4K

    10 Sources, first seen 10h ago

    Combined views

    252.4K

    10 Sources, first seen 10h ago

    2.2K likes

    Useful links

    arXiv.org

    Superhuman AI for Stratego Using Self-Play Reinforcement Learning...

    Summarized Science · YouTube

    New AI Conquers Deception, Crushes World's Best Player For Just $8,000
    10h ago
    first seen 10h ago
    2.2K likes
    62 comments
    1.2K saves
    184 reposts

    Ataraxos, an AI built to play Stratego, won 15 games against champion Pim Niemeijer in a 20-game series, with one loss and four draws, researchers report in a Nature paper published Sept. 30.

    Stratego makes the opponent’s piece identities a central uncertainty. Each player secretly arranges 40 pieces, concealing their identities from the other player. Battles reveal the pieces involved. Ataraxos combines learning through games against itself with a network that models this hidden information and search that evaluates possible moves, the paper explains.

    Playing with hidden pieces

    The researchers trained separate, connected networks for arranging pieces and choosing moves. A belief network predicts the types of the opponent’s concealed pieces. Before a move, the system samples possible board states and simulates candidate actions to estimate their value.

    The match series lasted three weeks, giving Niemeijer time to prepare between games and look for weaknesses. He was told that Ataraxos would not adapt to his play, according to the paper.

    The reported 85% effective win rate treats each draw as half a win. Ataraxos won 15 of the 20 games outright, so the effective metric should not be read as an 85% share of victories.

    What the result establishes

    The authors describe the performance as, to their knowledge, Stratego’s first superhuman result. That is their assessment of the evaluation, rather than a direct comparison against every previous system.

    They sought a match against DeepMind’s DeepNash, but report that DeepMind declined because its code was no longer functional. The paper therefore offers no head-to-head result between the two AIs.

    Noam BrownOpenAI
    Featured Source
    62 comments
    1.2K saves
    184 reposts

    Sentiment

    Positive74.4%25.6%Negative

    Summary

    Useful Links

    arXiv.org

    Superhuman AI for Stratego Using Self-Play Reinforcement Learning...

    Related Videos

    Sentiment

    Positive74.4%25.6%Negative

    Many accounts congratulated the team on the first superhuman Stratego AI after two years of work, while some replies questioned the result's novelty or the collaboration story.

    Based on 49 sentiment-bearing replies from 48 accounts across 5 conversations.

    New AI Conquers Deception, Crushes World's Best Player For Just $8,000Summarized Science · YouTube
    Today's Rank

    #11

    Today's Rank

    #11

    Summary

    Many accounts congratulated the team on the first superhuman Stratego AI after two years of work, while some replies questioned the result's novelty or the collaboration story.

    Based on 49 sentiment-bearing replies from 48 accounts across 5 conversations.

    Useful Links

    arXiv.org

    Superhuman AI for Stratego Using Self-Play Reinforcement Learning...

    Related Videos

    • New AI Conquers Deception, Crushes World's Best Player For Just $8,000Summarized Science · YouTube

    11 Sources

    NatureScalable decision-making for games of imperfect information
    @NatureNature research paper: Scalable decision-making for games of imperfect information https://go.nature.com/3TvlQQS
    @ssokotaIn our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
    @zicokolterRT @ssokota: In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed…
    @EugeneVinitskyTwo years of work and algorithmic innovation later, we have a clearly superhuman Stratego AI and the techniques scale generally to other imperfect information settings!
    @polynoamialI met @ssokota 5 years ago at a conference poster session. I was impressed with his explanation, so I offered him an internship. We continued collaborating and he now works with me at @OpenAI. Later, he told me I was the only person to stop by his poster for the entire session.
    @teortaxesTexbig result in RL for "human-dominant" domains with imperfect information. There are no human-dominant domains in the long run.

    Related

    OpenAI Researcher Explains Selective Math Result Announcements

    Noam Brown says OpenAI announces internal model math results only when they shift views on AI progress.

    Codex Engineer Shares Extended Context Setup for GPT-5.6 Sol

    OpenAI Codex engineer Tibo Sottiaux posted configuration steps to unlock the larger window for ChatGPT accounts.

    OpenAI Shares Astra Proofs for Ten Math Advances

    Internal Astra model produces Lean-certified proofs for ten open problems across math and theoretical computer science.

    OpenAI Astra Model Solves Ten Open Problems

    11 Sources

    NatureScalable decision-making for games of imperfect information
    @NatureNature research paper: Scalable decision-making for games of imperfect information https://go.nature.com/3TvlQQS
    @ssokotaIn our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed for RL & test-time compute under imperfect information. 1/N
    @zicokolterRT @ssokota: In our Nature paper, we introduce the first superhuman Stratego AI, which we built using general techniques that we developed…
    @EugeneVinitskyTwo years of work and algorithmic innovation later, we have a clearly superhuman Stratego AI and the techniques scale generally to other imperfect information settings!
    @polynoamialI met @ssokota 5 years ago at a conference poster session. I was impressed with his explanation, so I offered him an internship. We continued collaborating and he now works with me at @OpenAI. Later, he told me I was the only person to stop by his poster for the entire session.
    @teortaxesTexbig result in RL for "human-dominant" domains with imperfect information. There are no human-dominant domains in the long run.