Backpropagation: how neural networks learned from their mistakes (1986)
In 1986 David Rumelhart, Geoffrey Hinton and Ronald Williams showed that multilayer neural networks trained with backpropagation learn their own internal representations. The method is still the standard way to train most neural networks.
On October 9, 1986, a four-page paper appeared in Nature (volume 323) by David Rumelhart and Ronald Williams of the Institute for Cognitive Science at the University of California, San Diego, and Geoffrey Hinton of Carnegie Mellon University. It described a learning procedure called backpropagation, in which internal “hidden” units that are neither inputs nor outputs learn to represent important features of a task. The algorithm repeatedly adjusts weights to minimise the difference between the network’s actual output and the desired one, letting neural networks learn their own inner representations.
The problem of the hidden layers
Rosenblatt’s perceptron (article) learned by adjusting weights from errors. In 1969, Minsky and Papert showed in Perceptrons that one-layer perceptrons could not represent some functions (including XOR), helping trigger the first AI winter (article). Multi-layer networks could in principle escape those limits, but there was no practical way to train units that sat between input and output. The authors of the Nature paper argued that the simpler perceptron convergence procedure could not create new features on its own.
Passing the error backwards
The method works by a forward pass through the network, a comparison with the target output, and a backward pass using the chain rule from calculus. The error measure is half the sum of squared deviations between actual and desired outputs; weights change in proportion to the gradient, which is gradient descent. You can explore the idea interactively in Geek of the Week: Gradient descent and see depth in action in Geek of the Week: Neural network playground.
In one example, the network learned mirror symmetry using only two hidden units after 1,425 passes through all 64 possible input vectors. In another, two isomorphic family trees (English and Italian names) were trained on 100 of 104 possible triples (person, relation, person). The authors noted an obvious drawback: the error surface can contain local minima, so gradient descent does not guarantee a global optimum. They also wrote: “The learning procedure, in its current form, is not a plausible model of learning in brains.”
Not the first, but the most influential
Others had developed related ideas before. Seppo Linnainmaa published the “reverse mode” of automatic differentiation in 1970 (Helsinki master’s thesis). Paul Werbos described training neural networks with backpropagation in his 1974 Harvard dissertation and applied it to multi-layer nets in 1982. The Nature authors themselves credit David Parker and Yann LeCun with independent variants. The 2024 Nobel physics background paper notes that Rumelhart, Hinton and Williams “reinvented” a scheme others had applied before; more important was proving that networks with a hidden layer can learn tasks impossible without one.
A fuller account appeared in the same year in Parallel Distributed Processing (MIT Press, 1986). Yann LeCun and colleagues later used backpropagation to train convolutional networks on handwritten zip codes (1989); from the mid-1990s, several US banks used such nets to read digits on cheques (article).
Why it matters
The ACM’s 2018 Turing Award citation states that backpropagation is “standard in most neural networks today.” AlexNet was trained with stochastic gradient descent, the same broad family of methods. Hinton shared the 2024 Nobel Prize in Physics with John Hopfield. The core idea from 1986, measure the error and nudge every weight a small step downhill, still underpins most training of machine learning models built on neural nets.