Gerald Tesauro

IBM researcher who built TD-Gammon, a 1992 neural network that learned backgammon by self-play.

Tesauro’s TD-Gammon used temporal-difference reinforcement learning to reach strong backgammon play without handcrafted rules. It showed that neural networks could improve through experience during the quiet years after the second AI winter.