A game-playing program called Ataraxos has become the first artificial player to beat the best humans at Stratego, and how it did so may matter more than the win. Built by researchers at Carnegie Mellon, MIT, New York University and Stanford, it defeated Pim Niemeijer, the game's most decorated champion, by fifteen wins to one with four draws across a twenty-game match. The work was published in Nature and reported by MIT News on September 30.
Stratego has long been a hard target. Unlike chess or Go, each side hides the identity of its pieces, so a player has to reason under deep uncertainty about what the opponent is holding. That makes it a classic test of imperfect-information games, the messy category that most real-world decisions actually belong to.
Thinking at move time
DeepMind cleared the first bar here in 2022 with DeepNash, a system that reached expert human level. Ataraxos goes past it, and on a budget that is hard to credit. The team trained it on sixteen H100 graphics chips for about a week, running some 160 million practice games, at a cost its authors put at a few thousand dollars. DeepNash, by comparison, learned from roughly 5.5 billion games at a reported cost of three to four and a half million.
The saving did not come from a bigger model or more data. It came from letting the system plan at the moment it moves, using a generative model to imagine how the hidden board might really be laid out before it commits. The authors call that decision-time planning the missing piece. In plainer terms, the program stops to think during the game rather than trying to memorise the right answer to every situation in advance.
Not just one game
The same method travelled. The researchers pointed Ataraxos at Barrage Stratego, a faster variant, and at two card games with very different shapes: Hanabi, where several players cooperate, and Dou dizhu, where two gang up on a third. It reached superhuman play in each. That generality is the part worth watching, because a trick that works on one board is a curiosity, while a method that handles bluffing, cooperation and shifting alliances starts to look like a tool.
There is a through-line to recent work here. AI, Claudius has covered systems that cracked a nine-loop physics problem and settled a maths question open since the 1980s for modest sums. Ataraxos belongs to the same shift. The old headline was that an AI could do the hard thing at all. The new one is that it can do the hard thing cheaply, and that is the version that changes who gets to try.
Sources
- i. www.nature.com
- ii. news.mit.edu
- iii. techxplore.com
Commentarii · 0