Finite-time analysis of the multiarmed bandit problem¶
Auer, P., Cesa-Bianchi, & Fischer, P. (2002). Finite-time analysis of the multiarmed bandit problem. Machine Learning, 47(2–3), 2-3.
Cited by¶
1 citation across 1 artifact.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Variation Strategies
- The bandit framework balances exploitation (continuing with the best-known option) against exploration (testing alternatives), a problem Auer, Cesa-Bianchi, and Fischer (2002) solved with the UCB algorithm by quantifying the optimal exploration premium as a logarithmic function of regret.
This sourceEstablishes the UCB algorithm and proves logarithmic regret bounds, providing the canonical formal treatment of the exploration–exploitation trade-off in sequential decision-making.
- The bandit framework balances exploitation (continuing with the best-known option) against exploration (testing alternatives), a problem Auer, Cesa-Bianchi, and Fischer (2002) solved with the UCB algorithm by quantifying the optimal exploration premium as a logarithmic function of regret.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:5907ff79c386 · see in the full table