Self-play

Training regime in which an agent improves by competing against copies or past versions of itself, generating diverse experience without human expert labels.

Self-play is a training strategy in reinforcement learning where an agent plays against itself or against earlier snapshots of its own policy. Each game produces new states and outcomes that serve as training data, so the system can improve without a fixed dataset of expert moves. Because both sides grow stronger together, self-play can explore strategies that human teachers might never demonstrate and can scale to domains with enormous state spaces.

Gerald Tesauro’s TD-Gammon used self-play with temporal-difference learning to reach near expert-level backgammon in the early 1990s. Decades later, AlphaGo combined self-play at far larger scale with deep neural networks and tree search to master Go. See the TD-Gammon article and the AlphaGo article.