Neural language models: Bengio's fight against the curse of dimensionality (2003)
In 2003 Yoshua Bengio and colleagues published A Neural Probabilistic Language Model in the Journal of Machine Learning Research, learning distributed word representations to escape the limits of huge n-gram tables.
In 2003, Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin published A Neural Probabilistic Language Model in the Journal of Machine Learning Research (volume 3). An earlier version appeared at NIPS 2000 (now NeurIPS). The model predicted the next word from a learned vector representation of the previous words, attacking the curse of dimensionality that makes huge language model tables of Markov chain n-grams impractical as vocabularies grow.
Distributed word representations
Instead of storing a separate parameter for every possible word sequence, the network learned word embeddings: dense vectors in which similar words sit near each other. A neural network combined these vectors to estimate probabilities over the next token. The idea foreshadowed modern LLMs, though 2003 hardware limited model size.
From n-grams to neural text
Classical Markov chain language models, including simple n-gram counts explored by Claude Shannon in 1948, assume that only the last few words matter. That keeps inference fast but cannot share statistical strength between related words like “cat” and “dog.” Bengio’s neural network approach reused structure across words through shared embeddings, a step toward the statistical NLP revolution of the 2000s.
Why it matters
The paper is widely cited as an early landmark in word embeddings and neural language models. Bengio later shared the 2018 Turing Award with Geoffrey Hinton and Yann LeCun. The work bridges backpropagation era nets (article) and the Transformer age (article), and it complements deep belief networks from the same decade.