Boltzmann machines: when neural networks borrowed statistical physics (1985)
In 1985 David Ackley, Geoffrey Hinton and Terrence Sejnowski published a learning algorithm for Boltzmann machines: stochastic binary units, an energy function from physics, and slow sampling that later methods would try to avoid.
In 1985, David Ackley, Geoffrey Hinton, and Terrence Sejnowski published A Learning Algorithm for Boltzmann Machines in Cognitive Science (volume 9, issue 1). They described a Boltzmann machine: a neural network of stochastic binary units whose collective state follows an energy function borrowed from statistical physics, with probabilities given by the Boltzmann distribution. The network could learn hidden representations, but training relied on slow sampling with Markov chains, a cost later generative adversarial networks were designed to avoid.
From Hopfield nets to hidden structure
The work built on John Hopfield’s associative memory networks from 1982, which store and retrieve patterns using an energy landscape. Hinton and Sejnowski had already linked perceptual inference to such ideas in a 1983 paper. The 1985 algorithm added hidden units so the network could discover structure in data rather than only complete partial patterns. The Nobel Committee cited Hinton’s Boltzmann machine when awarding him the 2024 Nobel Prize in Physics with Hopfield.
Stochastic units and slow learning
Each unit flips between on and off states with probabilities that depend on its inputs and the global energy. Learning adjusts connection strengths so training patterns become likely states. In practice, estimating the statistics needed for learning requires lengthy MCMC sampling based on Markov chains: alternating between clamping visible units to data and letting the network settle, which made full Boltzmann machines hard to scale.
Restricted machines and contrastive divergence
Paul Smolensky described a related harmonium model in 1986, a precursor to the restricted Boltzmann machine (RBM), which forbids connections within a layer so inference is more tractable. For years RBMs remained difficult to train until Hinton introduced contrastive divergence in 2002, a faster approximation that made RBMs practical building blocks. Stacked RBMs later became the foundation of deep belief networks (article).
Why it matters
Boltzmann machines showed that unsupervised learning could extract features through probabilistic energy models, not only through supervised backpropagation. The sampling bottleneck also motivated alternatives such as unrolled inference networks and GANs. The idea that learning systems borrow physics formalisms, and that generation can be painfully slow without shortcuts, still shapes deep learning today.