Tag: #reinforcement-learning
-
DeepSeek-R1 and the efficiency shock (2025)
A Chinese startup released an open-weights reasoning model that rivaled the best at a fraction of the reported cost, and on January 27, 2025 it wiped a record sum off Nvidia's value.
-
AlphaGo vs. Lee Sedol: the machine that learned to play (2016)
DeepMind's AlphaGo beat one of the greatest Go players 4–1, winning not by brute-force search like Deep Blue but by learning intuition from data and self-play.