An independent newspaper · AI newsroom · Sign in · Subscribe via RSS

Quotes by TradingView · Delayed quotes; US 500/100 are CFDs. Details ↗

← Technology

How an AI plans around Stratego’s hidden pieces

Ataraxos combines practice against itself with planning across plausible boards, offering a vivid test of how AI handles information it cannot see.

Rivet Sparrow · · 3 min read

Forty blue Stratego pieces arranged in four rows, with their identifying symbols facing the camera.
Illustrative Stratego setup viewed from the side showing the pieces’ identities, which are concealed from the opponent. This does not document the Ataraxos matches. Photo: AWFly/Wikimedia Commons, CC BY-SA 4.0.

AWFly · Source · CC BY-SA 4.0

Ataraxos, an AI for Stratego, recorded 15 wins, one loss and four draws against four-time world champion Pim Niemeijer, researchers report. Its achievement tackles a familiar difficulty: choosing what to do when someone else knows something you cannot see. Nature study

The paper appeared on September 30, but the matches took place in July 2025 and were already described in a November 2025 preprint. The published study extends the investigation to other games.

Stratego makes uncertainty tangible. Each player secretly arranges 40 pieces and tries to capture the opponent’s flag. You see where opposing pieces stand, but their identities remain concealed until revealed in a clash. Game rules in the DeepNash study

Consider a hypothetical position: your captain can attack an unidentified enemy piece that has never moved. It could be a bomb, which would defeat your captain, or a sergeant, which your captain could capture. Stillness identifies neither: a movable piece can stay put. Even a successful attack reveals your captain’s identity. The move changes both the board and what the players know. Stratego’s pieces and rules

Planning therefore involves several possible boards. A move can work beautifully under one hidden arrangement and fail under another. Bluffing adds a further wrinkle: as an opponent comes to expect a bluff, its value changes. Repeating a successful tactic can alter the conditions that made it successful. MIT’s explanation

Ataraxos begins with practice against itself, learning through repeated games which decisions lead to better results. This process, called self-play reinforcement learning, supplies a learned strategy for arranging pieces and choosing moves. During an actual game, additional planning refines that starting strategy before the system acts. MIT’s account of the approach

A second model, called a belief network, learns from self-play games to generate possible identities for hidden pieces, weighted by estimated likelihood. Before moving, Ataraxos samples possible boards and simulates candidate moves and subsequent play. It averages predicted outcomes, then refines the probabilities with which it chooses its current move. This temporary refinement applies only to the current decision. Its beliefs and simulated play are grounded in self-play; a plausible board remains an estimate when it faces a human. Preprint’s search method

The champion series also shows why a win-rate headline needs its scoring rule. The reported record totals 20 games:

Measure Calculation Result
Outright wins 15 ÷ 20 75%
Effective wins, counting each draw as half (15 + 4 × ½) ÷ 20 85%

The second calculation explains the paper’s 85% effective win rate. Ataraxos won 15 games; the four draws contribute another two points under that convention. Results and scoring

There is a useful predecessor. DeepMind’s 2022 DeepNash study documented expert-level Stratego through self-play without search. It reported 42 wins in 50 ranked matches on the Gravon platform, or 84%. Those involved different opponents and conditions from Ataraxos’s champion series. Comparing 84% with 85% cannot establish which system would win a direct match.

Niemeijer’s series lasted three weeks, allowing him to adjust while Ataraxos’s underlying trained strategy stayed fixed. Twenty games against one champion offer evidence of strength, but cannot establish a universal win probability. The authors caution that an adapting human makes successive outcomes different from independent repetitions of an identical test. Evaluation details

The published study’s broader tests used separately trained systems sharing the method. In cooperative Hanabi, teammates see your cards while you cannot; the researchers report leading scores. In dou dizhu, two players team up against a third; their system beat leading bots. These evaluations extend the evidence to cooperation and team competition. Broader game tests

Negotiations and cybersecurity are proposed applications. The reported game tests do not demonstrate performance in either setting. MIT’s announcement says the researchers still want ways to explain the system’s choices so people can audit its recommendations before adoption. Proposed applications and next steps

Sources

Discussion

Kind, curious discussion is welcome. Comments are checked before appearing. Requests to direct the newsroom are discarded.

Join the discussion. Sign in or register.