TD-Gammon: learning backgammon by playing itself (1992)

Gerald Tesauro at IBM built TD-Gammon, a neural network that learned strong backgammon through temporal-difference reinforcement learning and self-play, reaching near expert-level play.

Event date: Published:
HistoryResearch

In the early 1990s, Gerald Tesauro at IBM built TD-Gammon, a neural network that learned to play backgammon at strong amateur and near expert level through reinforcement learning and self-play. The program used temporal-difference learning to update its evaluation of board positions after each move, without labelled expert moves for every state. Tesauro described practical training issues in a 1992 Machine Learning paper and presented the full TD-Gammon story in Communications of the ACM in March 1995.

Learning from its own games

Unlike supervised learning on a fixed dataset, TD-Gammon generated training signal by playing games against itself. After each move, temporal-difference learning nudged the network toward valuations that better predicted the final outcome. The approach scaled to the huge state space of backgammon in ways hand-coded rules struggled to match during the second AI winter (article).

A bridge to modern game AI

TD-Gammon was not the first program to learn a board game, but it was a vivid public example that neural networks plus trial-and-error learning could reach high skill. The same broad family of methods later powered AlphaGo (article), which combined reinforcement learning and self-play at far larger scale. Modern LLM tuning via RLHF also optimises behaviour from feedback, though the setup differs from board-game self-play.

Why it matters

TD-Gammon showed that learning systems could improve through interaction rather than only from static labels, a lesson that outlived the symbolic expert systems boom. It remains a reference point when explaining how machine learning survived the 1987–1993 downturn alongside Yann LeCun’s convolutional nets (article).

#neural-network#reinforcement-learning