MCTS was introduced by Rémi Coulom in 2006 ("Efficient Selectivity and Backup Operators in Monte-Carlo Tree Search") for the game of Go, where minimax was failing because the branching factor (~250) and game length (~150) made traditional alpha-beta search hopeless. MCTS dominated computer Go from 2007 to 2015 and reached strong-amateur-level play.
Silver et al. 2016 "Mastering the Game of Go with Deep Neural Networks and Tree Search" (Nature) -- AlphaGo combined MCTS with two neural networks (a policy net trained on 30 million human moves, then improved by RL self-play, plus a value net) and beat Lee Sedol 4-1.
Silver et al. 2017 "Mastering the Game of Go without Human Knowledge" (Nature) -- AlphaGo Zero dropped the human-game pretraining entirely; pure self-play with one combined policy + value network beat the original AlphaGo 100-0 in 40 days.
Silver et al. 2018 "A General Reinforcement Learning Algorithm that Masters Chess, Shogi, and Go through Self-Play" (Science) -- AlphaZero generalized the same algorithm; in 4 hours of TPU training surpassed Stockfish at chess.
Schrittwieser et al. 2020 "Mastering Atari, Go, Chess and Shogi by Planning with a Learned Model" (Nature) -- MuZero did all of the above without being given the rules of the game -- it learns a dynamics model from raw experience.
2024: DeepMind's AlphaProof + AlphaGeometry 2 reached IMO silver-medal level on the 2024 International Mathematical Olympiad, applying the same recipe to formal math.