← The collection

Explore or Exploit.

Trying an uncertain option can teach you something, even when another seems better now.

Interactive experimentintuitiveField note ·
Preparing the experiment…
THE SHORT VERSION

Explore or Exploit, explained.

The multi-armed bandit problem describes the tradeoff between exploration, which discovers better options, and exploitation, which uses the best option known so far.

01 / THE MECHANISM

Why it happens

Early results are noisy. An option that wins once may be weak, while a promising option may initially lose. Exploring buys information at the cost of occasionally choosing a lower-paying option. The value of that information depends on how many decisions remain.

A bandit problem asks how to balance immediate reward and learning from experiments.

Read the result

Change exploration and compare cumulative reward over the same horizon. The oracle benchmark knows the best machine in advance; it shows the cost of learning, not an achievable strategy for someone without that knowledge.

02 / FOLLOW IT THROUGH

A worked example

Choosing a newsletter subject line

  1. Three subject lines have unknown response rates. You can test them over repeated sends.

  2. Send some messages using alternatives while giving most traffic to the current leader.

  3. The tests can discover a better line, but their cost matters more when only a few sends remain.

OPTIONAL DEEPER DETAILGo deeper: inside the model

Inside this model

Three seeded machines have fixed hidden win chances. The comparison agent uses epsilon-greedy choice over sample means, beginning with one trial of each. The oracle always chooses the best hidden chance. The plot shows cumulative reward over 60 rounds.

03 / BEYOND THE EXPERIMENT

Where this idea is useful

A practical use

Trying a new page design while still showing the strongest current design.

CHECK YOUR INTUITION

A common misconception

THE TEMPTING CONCLUSION

“The first winner deserves all future choices.”

THE MORE USEFUL DISTINCTION

A small sample can misidentify the best option. Continued exploration helps correct that mistake, especially over a long horizon.

What this explanation leaves out

  • Independent stationary rewards and a single epsilon rule omit changing audiences, costs and delayed outcomes.
ONE MORE QUESTION

Does this experiment use Thompson sampling?

The automated comparison uses an exploration-based policy, not Thompson sampling. Several algorithms solve bandit problems using different ways to represent uncertainty.

TAKE THE IDEA WITH YOU

How much future opportunity remains to benefit from what you learn today?

Associated thinkers

Further reading

Explore the original research or the teaching reference behind this experiment.