Build1 publisher3 min readPublished
A Stratego AI trained for under $8,000 beat the game's most decorated player 15-1-4
Ataraxos, a Stratego AI from Carnegie Mellon, NYU, Stanford and MIT, beat Pim Niemeijer 15-1-4 after a training run that cost under $8,000. The savings came from a custom simulator and sample efficiency, so the price holds only for problems that simulate as cheaply.
The Engineer · Build desk

What happened
- The training took one week on 16 Nvidia H100 GPUs, plus four more days on four GPUs for the belief network that guesses hidden pieces.
- DeepMind's DeepNash trained on 1,024 TPU nodes for two to three months, a run the Ataraxos authors estimate at $3 million to $4.5 million at 2025 prices.
- The two systems have never played each other, because DeepMind told the researchers the DeepNash code no longer works.
- Ataraxos learned only by playing itself, under a regularization term that forces it to vary its setups and moves early in training.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- cost A replication budget has to cover building a fast simulator for the target problem, because the authors credit their custom GPU simulator for part of the saving and the $8,000 counts only GPU time.
- constraint Problems that cannot be written as a closed game with fixed rules cannot use this recipe, since every training signal Ataraxos got came from games against itself.
- precedent Earlier work skipped per-decision search as too hard; Ataraxos now shows it working at superhuman level, so later hidden-information agents can start from a tested design.
Those two runs add up to 3,072 GPU-hours: 2,688 on the H100s and 384 for the belief network [1]. Spread over those hours, a bill under $8,000 comes to less than about $2.60 per GPU-hour [2]. The authors priced the hours at 2025 rates, so the figure is GPU time [8]. Sokota wrote on X that the team was limited to academic computing resources and picked up plenty of tricks over the years [11].
The authors name two sources of the saving. One is a custom GPU simulator they wrote. The other is much higher sample efficiency [10]. By their count, Ataraxos needed about 1/30th of DeepNash's self-play games and 1/100th of its training examples, for roughly 1/500th of the compute cost [9]. Measured against their own DeepNash estimate, a run under $8,000 costs no more than 1/375th to 1/560th as much [3]. That estimate rests on a training schedule a DeepNash author recalled, according to the paper [7].
Hidden information is what made Stratego expensive. Each side sets up 40 pieces face down, and the paper counts more than 10^33 possible setups [5]. Texas Hold'em has 1,326 possible starting hands [6]. Sokota wrote that methods derived from poker AI worked only when there was little hidden information, and the paper says their computational effort grows with the amount of it [6].
The training schedule is the part I would copy first, because it costs no hardware. Strong Stratego needs randomness, the authors say, because predictable players get exploited [15]. Ataraxos begins with heavy pressure to vary its play and large learning steps. The researchers loosen that pressure as training goes on and shrink the steps to match, so late training makes only small corrections [16]. Without that control, learning in imperfect-information games tends to go in circles or turn chaotic, according to the paper [16]. The researchers compare the regularization to an energy reserve. Burn it too fast and the agent makes quick early progress, then loses the ability to keep learning and often becomes easy to exploit [17].
Ataraxos also spends compute during play. Before each move it uses the belief network to generate possible game states, plays out candidate moves and runs an extra learning step for that one decision [18]. The authors say earlier work skipped this kind of search because it was considered too hard [19]. The under-$8,000 figure covers the training runs [8], and the report does not put a price on the per-move compute.
Niemeijer is the one opponent both systems have faced on the record [2] [13]. At the 2023 world championship, DeepNash won 19 of 28 games but lost to most of the top players, Niemeijer included [13]. He has four world championships, 15 Dutch national titles and two online world championships, and spent more than 600 weeks at the top of the world rankings [3]. George Franka, the only player who has competed in every world championship since 1997, said Niemeijer is "the best Stratego player of all time" [4]. Counting each draw as half a point, Ataraxos scored 17 of 20 against him [4].
The authors describe Ataraxos as, to their knowledge, the first AI to reach superhuman play in Stratego [1]. The published evidence for that is one 20-game series against one player [2].
What to watch
- Release of the custom GPU simulator and training code, so the sub-$8,000 run can be reproduced outside the authors' lab.
- Matches against other top players beyond the single 20-game series with Niemeijer.
- An attempt to apply the decaying regularization schedule and per-move learning to a hidden-information problem that lacks a cheap exact simulator.