Word embedding

Dense vector representation of a word learned from co-occurrence or prediction context, capturing semantic and syntactic regularities in continuous space.

A word embedding is a learned dense vector that represents a word’s meaning and usage in context. Unlike sparse one-hot encodings or huge Markov chain n-gram tables, embeddings place related words near one another in a continuous space, so the model can generalise across similar terms. Yoshua Bengio and colleagues showed in 2003 that a neural network language model could learn such distributed representations while predicting the next word, attacking the curse of dimensionality that plagued earlier count-based models.

Later work such as word2vec made training embeddings at scale practical, and the idea became a foundation for modern natural language processing and today’s large language models. Embeddings also power retrieval systems that match queries to documents in vector space. See the neural language model article and timeline entry.