Skip to content

Science2 publishers2 min readPublished

Self-play system Ataraxos outplays top-ranked human Stratego players on a smaller training budget

Ataraxos, an AI from MIT, Carnegie Mellon, NYU and Stanford researchers, beat top-ranked human Stratego players by a large margin, MIT reports. The team says it was far cheaper to train than rival systems and also beat top players at other strategy games.

The Scientist · Science desk

Illustration accompanying Self-play system Ataraxos outplays top-ranked human Stratego players on a smaller training budget

What happened

  • Stratego keeps piece ranks hidden, and MIT puts the number of possible piece configurations above 10^66, far more than in chess.
  • Earlier efforts, including Google DeepMind's, spent millions of dollars on training and still could not beat top human Stratego players, according to MIT.
  • Ataraxos learned a 'blueprint strategy' by playing against itself many times, a method called self-play reinforcement learning.
  • Carnegie Mellon graduate student Samuel Sokota led the work and MIT's Gabriele Farina was senior author, with the paper appearing in Nature.

Compiled by The ScientistSomething wrong?How this is made

Why it matters

  • cost If the training-cost claim holds, strong hidden-information play no longer needs a corporate-scale budget, and university groups like this one can compete on the benchmark.
  • capability Wins at other strategy games with different rules suggest the method is not tuned to Stratego alone, and any use outside games would depend on that.
  • constraint The uses MIT suggests in negotiation, cybersecurity and military planning are projections; every demonstrated win so far comes from games whose rules both sides know in advance.

"The more you bluff, the more your opponent expects it, and the less each bluff is worth. It's not obvious how to reason about that," Sokota said [12]. He compared it with chess: "It's very different from a setting like chess, where the best move is still the best move no matter how often you've played it" [12].

In chess, a good move stays good. In Stratego, a move's value depends on what the opponent has learned from your earlier play. Farina said methods built for earlier games do not carry over. "With Stratego, there is an explosion of possible universes you might have to deal with. AI techniques that were developed for games like poker definitely could not scale in this setting," he said [13].

According to MIT, the team combined efficient training algorithms with new techniques built for decision-making when information is hidden [4]. The result, which they named Ataraxos [2], is a system that MIT says plays Stratego better than the next-best models while needing far less computing to train [3].

Those are MIT's descriptions [1][3]. The release does not report the figures needed to check them: how many games Ataraxos played against people, its win rate, who the top-ranked opponents were, or the training cost in dollars or compute. A win over humans is only as informative as the number of games and the strength of the players behind it.

The cost comparison is also specifically about training [3]. Training is a one-time bill. A system used on live decisions pays its running cost on every move, and a model that was cheap to train is not automatically cheap to run.

Farina's own summary goes further than a leaderboard result. "In the kind of imperfect information tasks you would face in reality, you often don't have the luxury of enumerating through all the possibilities. There are just too many. Having AI algorithms that are general purpose and can provably perform this challenging task so well is a big step forward," he said [14].

In my view, "provably" is the word in that statement most worth checking against the paper [10]. Beating ranked players measures a system against the people who sat down to play. A formal guarantee, depending on exactly what it covers, would also say something about opponents who never did.

What to watch

  • Independent matches, or a public release of Ataraxos, that let Stratego players and other labs test it outside the research team's own evaluation.
  • Replication of the cheaper training method by another lab, which would show whether the cost advantage holds away from its authors' hardware and tuning.
Loading claim ledger
Loading source directory links
Loading share composer
Loading topic controls
Loading related stories