Ataraxos reaches superhuman play in Stratego
On 30 September 2026, Nature published results for Ataraxos, a reinforcement learning and search system built by researchers at Carnegie Mellon, MIT, Stanford and NYU. In a 20-game official series the system defeated Pim Niemeijer — four-time world champion and, in the authors' description, the most decorated human Stratego player — by 15 wins to 1 loss with 4 draws, a result significant at P < 2.6 x 10^-4. Across a wider championship test the system finished 39-2. The authors describe it as the first superhuman result in the game's history. Training used 16 Nvidia H100 GPUs for about one week, roughly $8,000 of compute, which the paper puts at about one five-hundredth of the compute, one-thirtieth of the self-play games and one-hundredth of the training examples used by DeepMind's earlier Stratego system DeepNash, estimated at $3–4.5 million. The same method produced superhuman Barrage Stratego play and state-of-the-art results on two-to-five-player Hanabi and dou dizhu.
Why It Mattered
Stratego has been the standing hard case for imperfect-information game AI: both players' pieces are hidden, and the paper puts the number of possible configurations above 10^66, which defeats the search techniques that solved chess and Go and strains the equilibrium-finding methods that solved poker. DeepNash in 2022 reached strong expert play but never beat the top human tier. Clearing that bar closes a line of research that ran for roughly a decade. The more consequential number is the cost. A four-university academic team matched and exceeded an industrial laboratory result for about $8,000, inverting the assumption that progress on hidden-information decision-making requires frontier-scale compute. That matters beyond games: hidden-information search is the formal skeleton of negotiation, auction design, security games and multi-agent coordination, and bringing it back inside an academic budget changes who can contribute. The breadth claim is the third element — the same recipe transferred to a cooperative game (Hanabi) and a three-player card game (dou dizhu), which the authors frame as a general design pattern rather than a Stratego-specific engineering effort. If that generality holds under replication, the 2030 account will record this less as a game milestone than as the point where efficient search under hidden information became a commodity technique.
Who Built It
Carnegie Mellon University, with MIT, Stanford and NYU
Applications
- Game Playing
- Reinforcement Learning
- Multi-Agent Systems