Gradient descent: watch a model roll downhill to learn

Drop a ball on a loss curve and watch it step downhill. Too large a learning rate and it flies off; a bumpy landscape traps it. This is how almost every neural network is trained.

y26w36

This week the machine does the one thing behind almost all modern AI training: it rolls downhill. The perceptron learned a line, but how does any model actually improve? It measures how wrong it is (the loss), works out which way is less wrong, and takes a small step that way. Repeat a few billion times and you have trained a network.

The gadget below is a gradient descent playground. The curve is a loss landscape: low points are good (small error), high points are bad. The ball is the model’s current setting. At each step it looks at the slope under its feet (the gradient) and steps downhill by an amount set by the learning rate.

Try this:

  1. Click anywhere on the curve to drop the ball, then press Run. With a small learning rate on the Bowl, it glides smoothly to the bottom.
  2. Now crank the learning rate up and Run again. Too big a step and the ball overshoots the valley, bounces higher each time, and diverges. This is the single most common reason training blows up.
  3. Switch to Hills and start the ball on one side. With a gentle rate it often settles into the nearest dip, a local minimum, not the deepest one. A larger step can jump it out toward a better valley.

Why this is (almost) all of training

Every weight in a neural network is one dimension of a landscape like this, except the real thing has millions or billions of dimensions instead of one. Training means measuring the slope in all of them at once (that is what backpropagation does) and taking a downhill step. The learning rate you just played with is a real, fiddly knob that every practitioner tunes.

This one idea, scaled up with fast hardware and lots of data, is what made deep learning work: it trained AlexNet to see and the transformer to read. Behind the magic is a ball, a slope, and a small step in the right direction.

Missed the earlier editions? Train a perceptron, watch a Markov chain babble, lose to unbeatable tic-tac-toe, teach a filter to see edges, watch an agent learn from reward, or see how a model reads text in tokens.

Click the curve to drop the ball, set a learning rate, and press Run. The ball follows the slope downhill. Too high a rate and it overshoots and flies off; a bumpy landscape can trap it in the wrong valley.