Search + Learning: The Idea Behind AlphaGo, AlphaZero, MuZero
This is the recipe that beat Lee Sedol, crushed Stockfish in 4 hours of training, and reached IMO silver-medal level on Olympiad math problems in 2024. MCTS guided by a neural network — the network gives the search a smart prior, the search gives the network a stronger target. They bootstrap each other into superhuman play. And in 2025 the same recipe shows up in o1, o3, and DeepSeek-R1's reasoning. Search + learning may be the most general AI idea we have.
Scope of this lesson — the IDEA. This page teaches MCTS (the general algorithm) and tells the AlphaGo → AlphaZero → MuZero story at a conceptual level: why search and learning compose, where each step in the progression removed an assumption, and how the recipe escapes board games. For the math — the PUCT derivation, the term-by-term AlphaZero loss, the MuZero K-step loss, and a full AlphaZero (not just MCTS) tic-tac-toe trainer — read the companion lesson AlphaZero & MuZero: Search + Learning in One Loop after this one.
After this lesson, you will be able to:
- Walk through MCTS step by step — selection (UCB1), expansion, simulation, backup — and understand why it became the planning algorithm of choice for combinatorial games
- Trace the AlphaGo → AlphaGo Zero → AlphaZero → MuZero progression: each step removes more human knowledge while keeping the same core recipe — deep network + search + self-play
- See how MCTS acts as a policy improvement operator at training time: the search produces a stronger policy than the network alone, and the network learns to imitate that stronger policy
- Recognize the recipe in modern systems beyond board games — AlphaProof's IMO-medal performance and o1/o3 reasoning models are the same idea applied to math and language
Before You Start
Don't worry if "search + learning" sounds magical — once you implement MCTS once, you'll see it's just bookkeeping over a tree. The deep network on top is what makes it scale, but the core algorithm is something you can write in a single afternoon.
#The Two Halves of the Recipe
#Monte Carlo Tree Search: The Four Phases
MCTS builds a search tree iteratively. Each iteration goes through four phases:
#AlphaGo: Networks + Search
The original 2016 AlphaGo had three neural networks:
- Supervised policy network p_σ trained on 30 million human moves — predicts the next human move given the board.
- RL policy network p_ρ — the supervised network further trained against itself via REINFORCE-style policy gradient.
- Value network v_θ — predicts the winner from a given position.
At play time, MCTS used p_σ as a prior on which moves to explore (multiplied into the UCB selection), called the value network at leaf nodes for evaluation, and combined it with rollouts using a fast linear policy.
The architecture worked but was complex. AlphaGo Zero stripped it down.
#AlphaGo Zero: Self-Play From Scratch
Two key changes in 2017:
- One combined policy + value network (two heads on a shared ResNet backbone) replaces the three networks.
- No human data. The network starts with random weights. It plays itself with MCTS, and the MCTS-improved move distributions become the training labels.
The training signal is a single loss with three terms: value MSE (predicted value vs. actual outcome), cross-entropy between the MCTS visit-count distribution and the network's policy, and L2 weight decay. The MCTS at each position runs ~1600 simulations; the visit-count distribution at the root is treated as the improved policy π, and the network is trained to match it.
For the math. The full AlphaZero loss term by term, the PUCT vs. UCB1 distinction, and the visit-count softmax with temperature τ all live in the companion lesson AlphaZero & MuZero: Search + Learning in One Loop. This page sticks to the idea.
#AlphaZero: One Algorithm, Three Games
AlphaZero (2018) applied AlphaGo Zero to chess and shogi as well as Go. Same code, same network architecture, same algorithm. After 4 hours of training on 5000 first-gen TPUs:
- Chess: surpassed Stockfish 8 (then-state-of-art).
- Shogi: surpassed Elmo (the strongest shogi engine).
- Go: surpassed AlphaGo Zero.
The lesson was profound: a single algorithm with no game-specific knowledge could master three completely different games at superhuman level. Decades of hand-engineered domain knowledge in chess engines were defeated by general search + general learning + self-play.
You want to apply AlphaZero-style training to a new game where you DON'T have a perfect simulator (you can't query 'if I take action a, what state results?'). What's your problem and what's the fix?
The answer is the bridge to MuZero: AlphaZero's MCTS expansion phase requires knowing how the world reacts to each action. For board games, you have the rules. For Atari games, you have the emulator. But for many real-world problems, you don't have a perfect simulator.
#MuZero: Learn the Rules Too
MuZero matched AlphaZero on board games and dramatically beat the prior state-of-art on Atari — the first algorithm to do both top-tier game playing and top-tier video-game-from-pixels with one method.
For the math. The three networks (h, g, f), the K-step unrolled loss that keeps the learned dynamics from collapsing, and a runnable AlphaZero-on-tic-tac-toe playground (not just MCTS) are in AlphaZero & MuZero: Search + Learning in One Loop.
Tests · Verify each move is legal. Verify game terminates within 9 turns. Verify visit counts at root sum to n_sims. Verify the move chosen is the most-visited child.
#The Modern Recipe
#Key Takeaways
- MCTS = selection (UCB) + expansion + simulation + backup, repeated thousands of times. The visit count distribution at the root is the search-improved policy. The action with highest visit count is the move.
- AlphaZero's loss is one elegant equation. Match the network's value to the game outcome (MSE), match the network's policy to MCTS visit counts (cross-entropy), L2 regularize. Train across millions of self-play positions.
- MCTS is a policy improvement operator. Search converts a policy in (the prior network) to a stronger policy out (the visit-count distribution). The network learns to imitate the stronger policy. The next iteration's search is even better.
- AlphaGo Zero → AlphaZero → MuZero progressively removed assumptions. AlphaGo Zero dropped human data; AlphaZero dropped game-specific tuning; MuZero dropped the requirement of perfect game rules by learning a dynamics model.
- The recipe generalizes far beyond games. AlphaProof, AlphaGeometry, and o1/o3-style reasoning models are non-game applications of search + learning + self-play. The main constraint: you need a verifiable reward signal to guide the search.
Interactive Lab
Step through Monte Carlo Tree Search one iteration at a time — see how the UCB1 bandit at every node grows the tree toward promising branches, and how visit counts (not raw Q-values) become the policy improvement signal AlphaZero learns from.
#Quick Check
Why does AlphaZero pick its move by visit count rather than by predicted value Q?