AlphaGo vs. Lee Sedol: the machine that learned to play (2016)

DeepMind's AlphaGo beat one of the greatest Go players 4–1, winning not by brute-force search like Deep Blue but by learning intuition from data and self-play.

In March 2016, in a Seoul hotel, DeepMind’s AlphaGo beat Lee Sedol, one of the strongest Go players of his generation, four games to one. It was watched live by tens of millions of people, especially across East Asia, where Go is treated as a serious intellectual art.

The result mattered because Go had been considered out of reach. The board is so large that the number of possible games dwarfs the number of atoms in the observable universe. You cannot win by searching every move, the way Deep Blue had done at chess in 1997.

How AlphaGo won

AlphaGo did not out-calculate Lee Sedol. It learned to judge positions the way strong players do, by feel, and then searched selectively around those judgements.

  • It first studied a large database of human games with deep learning, building neural networks that could predict good moves and estimate who was ahead.
  • It then improved by reinforcement learning, playing millions of games against versions of itself and keeping what worked.
  • Finally it combined those learned instincts with a focused tree search, spending its calculation only on lines the networks thought were promising.

The moves people remember

Two moments became famous. In Game 2, AlphaGo played move 37, a shoulder hit so unusual that professional commentators assumed it was a mistake, until it slowly proved decisive. Then in Game 4, Lee Sedol answered with move 78, a wedge so precise that AlphaGo lost its footing and, eventually, the game. It was the only game a human would win in the match, and it is remembered as a moment of human brilliance under pressure.

Why it matters

Where Deep Blue had searched, AlphaGo learned. That is the whole difference. A brittle, hand-tuned evaluation function was replaced by intuition distilled from data and self-play, exactly the recipe that AlexNet had shown could work for vision, now applied to strategy.

The approach generalised. Its successor, AlphaGo Zero, learned Go from scratch with no human games at all, and later systems extended the same self-play idea to chess, shogi, and beyond. AlphaGo did not just win a board game; it showed that learned intuition could beat human experts at tasks long thought to require something uniquely human.

#reinforcement-learning #deep-learning