Multi-Agent RL
AlphaGo beat Lee Sedol via self-play. Pluribus beat the top human poker pros via CFR. AlphaStar reached StarCraft Grandmaster via a league of agents. Cicero negotiated alliances in Diplomacy at human level. Every superhuman game-playing AI is multi-agent, and the moment you have two learners in the same room, the convergence proofs you grew up with stop working. This is where RL gets strange and beautiful.
After this lesson, you will be able to:
- Distinguish cooperative, competitive, and mixed multi-agent settings, and pick the right algorithm family for each
- Use self-play to train two-player zero-sum agents that converge to Nash equilibria — the recipe behind AlphaGo, AlphaZero, AlphaStar
- Apply Counterfactual Regret Minimization (CFR) to imperfect-information games like poker — the math that powers Pluribus and Cepheus
- Pick between centralized-critic (MADDPG / QMIX / MAPPO) and independent learners, and know why naive PPO often surprisingly works
Before You Start
Don't worry if "multi-agent" sounds intimidating — once you have one PPO agent working, multi-agent is mostly variations on "let multiple PPO agents play each other and see what falls out". The hard parts are the math of equilibria and the tricks for stability.