Language model

Statistical or neural model that assigns probabilities to word sequences, used for prediction, generation, and as a backbone for modern NLP systems.

A language model estimates the probability of the next word given preceding context, or more generally the likelihood of a text sequence. Early models used Markov chain n-grams: tables of word counts that grow explosively as vocabulary and context length increase. Claude Shannon demonstrated simple Markov text generation in 1948, but practical systems long struggled with sparse data and limited generalisation.

Yoshua Bengio and colleagues proposed a neural network language model in 2003 that learned word embeddings jointly with next-word prediction, sharing statistical strength across related terms. That line of work led to recurrent and transformer architectures and ultimately to today’s large language models, which scale the same basic objective to billions of parameters. See the neural language model article and timeline entry.