Ataraxos AI beat a Stratego champion 15–1, MIT researchers report
The system combines self-play with a search for plausible hidden pieces. Its reported wins show progress in a difficult game, while proposed real-world uses remain untested.
Ataraxos, an AI system developed by researchers at MIT, Carnegie Mellon University, New York University and Stanford University, beat four-time Stratego world champion Pim Niemeijer in a 20-game series played in July 2025, according to the researchers’ paper. MIT reported the result on September 30, 2026. The system won 15 games, lost one and drew four, a result that advances AI play in a game where each opponent’s pieces begin hidden.
The result is a test of strategic play under incomplete information. It does not show that Ataraxos can improve the military planning, business negotiations or cybersecurity decisions that MIT names as possible future applications. Those settings have consequences and constraints that the reported game evaluations did not test.
How Ataraxos plays Stratego
In Stratego, each player privately arranges 40 pieces and tries to capture the other player’s flag. Piece identities remain concealed until encounters reveal them. A player must therefore choose moves while reasoning about positions that might be true, rather than a fully visible board. The researchers describe the number of possible arrangements as a central obstacle for AI methods that try to account for every hidden possibility.
The researchers trained Ataraxos through two linked self-play processes. One learns how to arrange pieces at the start; the other learns how to move them during play. The processes affect one another because the starting layouts shape the games used to train move selection, and the outcomes of those games inform both processes. This separates a decision made before play begins from the sequence of decisions that follows.
Ataraxos also refines its choices while a game is under way. The paper describes a belief network that samples plausible identities for concealed opposing pieces, then evaluates candidate moves through simulated continuations. That gives the system a way to concentrate on likely versions of the current position when choosing its next move. Its training and in-game planning both contribute to the approach described by the team.
The paper says the final training run began in late May 2025 and used 16 Nvidia H100 graphics processors for a week. Training the belief network then used four H100s for four days. Those details describe the researchers’ reported computing setup; they do not establish what another group would spend to reproduce the result under different conditions.
What the matches show
The Niemeijer series ran over three weeks. The paper says the champion could adjust his play between games, while Ataraxos could not adapt between games in that way. It also discloses that Niemeijer was paid to participate and received bonuses for wins and draws. Those conditions matter when interpreting the 15-win result: it was a specified contest against an expert player, rather than a general test of every possible opponent or rule setting.
The researchers also report a demonstration at the Stratego World Championship on August 1–3, 2025. Their accessible manuscript records 38 wins and two losses in 40 games against attendees. MIT’s September announcement gives a different figure, 39–2, and describes the opponents as top players. The accounts differ in both total and description, so the manuscript’s specified 40-game record should not be treated as confirmation of MIT’s larger tally.
MIT says the team adapted Ataraxos to Barrage Stratego, Hanabi and Dou dizhu, which involve different rules and forms of hidden information, and reports strong performance in each. Those additional game results suggest the method is not confined to the standard Stratego board. They still measure play in games, not whether the system provides sound recommendations to people facing real-world decisions.
How it compares with DeepNash
Ataraxos follows an earlier Stratego milestone from Google DeepMind. In December 2022, DeepMind reported that its DeepNash system learned through self-play without in-game search. It reached an all-time top-three ranking on the Gravon Stratego platform and won 84% of its games against expert players there. DeepMind’s researchers pointed to hidden pieces and games that can last hundreds of moves as reasons techniques from chess, Go and poker were difficult to apply.
MIT says Ataraxos achieved greater playing strength than DeepNash while using fewer than one hundredth as many training examples and fewer than one thirtieth as many self-play games. That comparison comes from the Ataraxos researchers’ evaluation. The reported human matches provide a separate measure of Ataraxos’s play, but the supplied records do not establish an independent, matched replication of its performance or cost advantage over DeepNash.
What remains before real-world use
MIT identifies military maneuvers, business negotiations and cybersecurity as possible applications of decision-making under hidden information. It lists partial support for the research from the Office of Naval Research, the National Science Foundation, NYU-related sources and a Schmidt Sciences fellowship. Neither the funding nor the game results amount to a demonstrated deployment in any of those settings.
Senior author Gabriele Farina says the team wants Ataraxos to explain its decisions so people can audit recommendations before adoption. ‘Humans must have the final say in whether a recommendation is followed,’ he told MIT News. For now, the documented advance is in how the system learns and plans in games with hidden information; the researchers have not shown that its choices are reliable or understandable enough for the proposed uses outside them.
Sources and context
- This game-playing AI is the new champ at StrategoMIT News
- Superhuman AI for Stratego Using Self-Play Reinforcement Learning and Test-Time SearcharXiv; Sokota and co-authors
- Mastering Stratego, the classic game of imperfect informationGoogle DeepMind
- Mastering the game of Stratego with model-free multiagent reinforcement learningScience; indexed by PubMed
AI-assisted article checked against the listed sources. NewsJaws did not conduct interviews or attend the reported events.
About NewsJaws Desk
AI-assisted reporting and explainers reviewed against the linked source documents. No claim of on-scene reporting or original interviews.