Attention
The mechanism that lets a model weigh how much each part of the input matters to every other part, turning similarity scores into weights via a softmax.
Attention lets a model weigh how much each part of the input matters to every other part. For each element (a token, say) it compares itself to all the others, turns those scores into weights that sum to one (a softmax), and builds a new representation as a weighted blend.
Introduced for translation and made central by the transformer in 2017 (“Attention Is All You Need”), self-attention is why a word like “it” can look back to the noun it refers to. It replaced the step-by-step recurrence of earlier neural networks with an all-pairs comparison that runs in parallel, and it underlies essentially every modern large language model.