Explore-Exploit Tradeoff
The core decision in reinforcement learning and foraging: spend a turn on the option you know pays off, or spend one gathering information that might pay off more.
Explore-Exploit Tradeoff
The multi-armed bandit problem asks: you are pulling levers on a row of slot machines with unknown, different payout rates, and every pull is a turn you cannot spend twice. Do you keep pulling the lever that has paid off so far, or spend a turn on a lever you know less about. John Gittins solved a clean version of this in 1979 with an index that ranks every option by the value of what pulling it would teach you, not just what it is expected to pay. Brian Christian and Tom Griffiths popularized the idea for a general audience in Algorithms to Live By (2016), including the related “look then leap” result: spend roughly the first third of your decision window exploring before you commit.
The tradeoff shows up any time a fixed budget of turns has to split between what already works and what might work better. Time, unlike money, cannot be borrowed against the future, which makes this the same problem underneath a research and development budget, a dating life, and a list of hobbies competing for your Saturday.