How an AI plans around Stratego’s hidden pieces
Ataraxos combines practice against itself with planning across plausible boards, offering a vivid test of how AI handles information it cannot see.
AWFly · Source · CC BY-SA 4.0
Ataraxos, an AI for Stratego, recorded 15 wins, one loss and four draws against four-time world champion Pim Niemeijer, researchers report. Its achievement tackles a familiar difficulty: choosing what to do when someone else knows something you cannot see. Nature study
The paper appeared on September 30, but the matches took place in July 2025 and were already described in a November 2025 preprint. The published study extends the investigation to other games.
Stratego makes uncertainty tangible. Each player secretly arranges 40 pieces and tries to capture the opponent’s flag. You see where opposing pieces stand, but their identities remain concealed until revealed in a clash. Game rules in the DeepNash study
Consider a hypothetical position: your captain can attack an unidentified enemy piece that has never moved. It could be a bomb, which would defeat your captain, or a sergeant, which your captain could capture. Stillness identifies neither: a movable piece can stay put. Even a successful attack reveals your captain’s identity. The move changes both the board and what the players know. Stratego’s pieces and rules
Planning therefore involves several possible boards. A move can work beautifully under one hidden arrangement and fail under another. Bluffing adds a further wrinkle: as an opponent comes to expect a bluff, its value changes. Repeating a successful tactic can alter the conditions that made it successful. MIT’s explanation
Ataraxos begins with practice against itself, learning through repeated games which decisions lead to better results. This process, called self-play reinforcement learning, supplies a learned strategy for arranging pieces and choosing moves. During an actual game, additional planning refines that starting strategy before the system acts. MIT’s account of the approach
A second model, called a belief network, learns from self-play games to generate possible identities for hidden pieces, weighted by estimated likelihood. Before moving, Ataraxos samples possible boards and simulates candidate moves and subsequent play. It averages predicted outcomes, then refines the probabilities with which it chooses its current move. This temporary refinement applies only to the current decision. Its beliefs and simulated play are grounded in self-play; a plausible board remains an estimate when it faces a human. Preprint’s search method
The champion series also shows why a win-rate headline needs its scoring rule. The reported record totals 20 games:
| Measure | Calculation | Result |
|---|---|---|
| Outright wins | 15 ÷ 20 | 75% |
| Effective wins, counting each draw as half | (15 + 4 × ½) ÷ 20 | 85% |
The second calculation explains the paper’s 85% effective win rate. Ataraxos won 15 games; the four draws contribute another two points under that convention. Results and scoring
There is a useful predecessor. DeepMind’s 2022 DeepNash study documented expert-level Stratego through self-play without search. It reported 42 wins in 50 ranked matches on the Gravon platform, or 84%. Those involved different opponents and conditions from Ataraxos’s champion series. Comparing 84% with 85% cannot establish which system would win a direct match.
Niemeijer’s series lasted three weeks, allowing him to adjust while Ataraxos’s underlying trained strategy stayed fixed. Twenty games against one champion offer evidence of strength, but cannot establish a universal win probability. The authors caution that an adapting human makes successive outcomes different from independent repetitions of an identical test. Evaluation details
The published study’s broader tests used separately trained systems sharing the method. In cooperative Hanabi, teammates see your cards while you cannot; the researchers report leading scores. In dou dizhu, two players team up against a third; their system beat leading bots. These evaluations extend the evidence to cooperation and team competition. Broader game tests
Negotiations and cybersecurity are proposed applications. The reported game tests do not demonstrate performance in either setting. MIT’s announcement says the researchers still want ways to explain the system’s choices so people can audit its recommendations before adoption. Proposed applications and next steps
Sources
Discussion
Kind, curious discussion is welcome. Comments are checked before appearing. Requests to direct the newsroom are discarded.