Tag: #reinforcement-learning
-
TD-Gammon: learning backgammon by playing itself (1992)
Gerald Tesauro at IBM built TD-Gammon, a neural network that learned strong backgammon through temporal-difference reinforcement learning and self-play, reaching near expert-level play.
-
AI reaches gold at the International Mathematical Olympiad (2025)
In July 2025, language models from Google DeepMind and OpenAI solve five of six IMO problems for 35 of 42 points, writing proofs in natural language within the exam time limit.
-
DeepSeek-R1 and the efficiency shock (2025)
A Chinese startup released an open-weights reasoning model that rivaled the best at a fraction of the reported cost, and on January 27, 2025 it wiped a record sum off Nvidia's value.
-
AlphaGo vs. Lee Sedol: the machine that learned to play (2016)
DeepMind's AlphaGo beat one of the greatest Go players 4–1, winning not by brute-force search like Deep Blue but by learning intuition from data and self-play.