Attention: watch words decide who to listen to

Click a word and watch arcs fan out to the others, thicker where it pays more attention, always adding up to 100 percent. Drag words together to make them attend. This is the mechanism behind every modern LLM.

y26w39

Last week a network learned to bend a boundary. This week we meet the idea that powers ChatGPT and almost every model like it, and it is surprisingly simple: let each word look at every other word and decide how much to listen to each one. That is attention, the beating heart of the transformer.

The gadget below is a sentence, one word per token. Click a word to make it the one looking around. Arcs fan out to every other word; the thicker the arc, the more attention it pays there, and the percentages always add up to 100. That last part is a softmax: a pile of raw scores squashed into a set of weights that sum to one.

Try this:

  1. Click the pronoun (“it”). See where its attention lands. Now drag “it” right next to “cat” and watch the arc to “cat” thicken: the pronoun learns to look back at the thing it stands for. This is exactly the trick that lets a model resolve what “it” means.
  2. Drag a word off on its own. Its attention spreads thin across everything. Pull it back into the crowd and the weights sharpen again.
  3. Slide Focus. Low focus spreads attention evenly across all words; high focus makes each word attend almost entirely to its nearest neighbor. Real transformers tune this same sharpness.

Why this matters

For decades, models read text one step at a time, passing a summary along like a game of telephone; distant words easily got lost. In 2017 a paper with the cheeky title “Attention Is All You Need” threw out the step-by-step part and kept only attention: every token compares itself to every other in one parallel sweep. That architecture, the transformer, is what made models scalable enough to become the chatbots and large language models we use today.

One honest note: a real transformer scores similarity with a dot product of learned vectors, not by closeness on a canvas. We use distance here so you can steer it by hand, but the shape of the idea is identical: score, softmax, blend. Behind the magic is a room full of words, each quietly deciding who to listen to.

Missed the earlier editions? Train a perceptron, watch a Markov chain babble, lose to unbeatable tic-tac-toe, teach a filter to see edges, watch an agent learn from reward, see how a model reads text in tokens, watch one roll downhill to learn, let a machine find groups on its own, or see a line learn to bend.

Attention is how a transformer lets each word look at the others. Click a word to select it: the arcs show how much of its attention goes to each other word (thicker means more), and the percentages always add up to 100. Drag words closer to make them attend to each other, and use Focus to sharpen or soften the attention.