Multi-armed bandit: explore or exploit?

Five slot machines hide different payout rates. Pull them yourself or slide Explore and watch epsilon-greedy learn which arm pays best. The classic reinforcement-learning dilemma in one row of levers.

y26w40

Last week words learned who to listen to. This week the machine faces a problem every recommender, ad system, and RL agent knows by heart: you have several options, you do not know which is best, and every try costs you something. Pull the wrong lever and you miss out. Never try new levers and you might never find the best one. That tension has a name: explore vs exploit.

The gadget below is a row of five one-armed bandits, old-school slot machines. Each pays on a hidden schedule. Some arms are duds; one pays better than the rest. You can pull by hand and feel your own curiosity, or slide Explore and let epsilon-greedy auto-play: with probability ε it tries a random arm (explore), otherwise it sticks to the arm with the best average so far (exploit).

Try this:

  1. Pull each arm a few times by hand. The bars show your running win rate; the numbers below are how often you pulled. Which arm looks best after ten pulls? Are you sure?
  2. Set Explore to 80% and hit Auto ×10. Watch the agent bounce around randomly, barely committing. Then drop Explore to 5% and auto-play again: it hammers the current favourite, maybe the wrong one.
  3. Hit Reveal odds to see faint ghost bars, the true payout rates. How much reward did you leave on the table by exploring too long, or by exploiting too early? That gap is regret, the currency of bandit algorithms.

Why this matters

The multi-armed bandit is the cleanest picture of reinforcement learning when there is no map to learn, only choices and feedback. News feeds, A/B tests, clinical trials, and game-playing AIs all face the same question: spend traffic on what already works, or gamble on something unknown that might be better?

Real systems use cleverer rules than epsilon-greedy (UCB, Thompson sampling), but the trade-off is identical. AlphaGo explored new moves in self-play; RLHF that tuned ChatGPT explores phrasing humans might prefer. Your row of five levers is the same loop at toy scale: act, observe reward, update beliefs, decide again.

Missed the earlier editions? Train a perceptron, watch a Markov chain babble, lose to unbeatable tic-tac-toe, teach a filter to see edges, watch an agent learn from reward, see how a model reads text in tokens, watch one roll downhill to learn, let a machine find groups on its own, see a line learn to bend, or watch words decide who to listen to.

Each arm pays out on a hidden schedule. Pull by hand to explore, or slide Explore to let epsilon-greedy auto-play: low = stick to what looks best, high = try random arms. Watch averages converge and regret shrink when you find the winner.