Temporal-difference learning (TD learning)

Reinforcement-learning method that updates value estimates from the difference between successive predictions, bootstrapping from experience without waiting for a final reward.

Temporal-difference learning (TD learning) is a family of algorithms in reinforcement learning that improve predictions by comparing successive estimates of future reward. Rather than waiting until a game or episode ends, the learner updates its value of the current state from the temporal difference between what it expected next and what it actually observed. This bootstrapping lets agents learn from partial experience in large state spaces where exhaustive labelling is impossible.

Richard Sutton formalised the approach in the 1980s, and Gerald Tesauro applied it to backgammon in TD-Gammon, a neural network that reached strong play through self-generated training signal. TD methods remain central to modern game-playing systems and value-based RL. See the TD-Gammon article and timeline entry.