# Contextual Bandits

> A contextual bandit is a learning algorithm that picks one action per context, sees only that action's reward, and balances exploration with exploitation.

Source: https://metavert.io/contextual-bandits  
Published: 2026-10-07  
Updated: 2026-10-07

A **contextual bandit** is a sequential decision algorithm that, on each round, observes some context (features of a user or situation), chooses one action from a set, and sees the reward for that action only. It then updates its policy so that future choices earn more. The name extends the multi-armed bandit, a row of slot machines with unknown payouts, by letting the best arm depend on who is pulling it. Contextual bandits are the standard tool for personalised choices with fast feedback: which article, offer, message or price variant to show to which person. They sit between a fixed A/B test, which learns one answer for everybody, and full [reinforcement learning](https://metavert.io/reinforcement-learning), which also models how today's action changes tomorrow's state.

### The Problem: Partial Feedback

In supervised learning every example comes with the right answer. A bandit sees only the outcome of what it chose; it never learns what would have happened under the alternatives. That creates the explore/exploit trade-off. A policy that always takes the action it currently believes is best never gathers evidence about the others, and can lock in an early mistake. A policy that explores too much wastes traffic on options it already knows are poor. Performance is measured as regret: the reward lost relative to always playing the best action for each context.

### Core Algorithms

**Epsilon-greedy** takes the best-looking action most of the time and a uniformly random one with small probability. It is crude but simple.

**LinUCB** (Li, Chu, Langford and Schapire, WWW 2010) models each action's expected reward as a linear function of the context and adds an upper-confidence bonus that is large where the model has seen little data, so uncertain options are tried because they might be good. In the original paper, on about 33 million events from a news-recommendation log, it achieved a 12.5% click lift over a context-free bandit, with a larger advantage when data was scarce.

**Thompson sampling** keeps a posterior distribution over each action's reward model, draws one sample from it, and acts greedily on the sample. Exploration comes from posterior uncertainty and fades as data accumulates. Chapelle and Li (NeurIPS 2011) showed it was competitive with the alternatives of the time and argued that it belonged among the standard baselines; Agrawal and Goyal (2012) later proved a regret bound of order d 3/2 √T for the linear case, within a √d factor of the lower bound.

Modern systems often replace the linear model with [boosted trees](https://metavert.io/xgboost) or neural networks and reduce the bandit to repeated supervised learning. An empirical "bake-off" by Bietti, Agarwal and Langford (JMLR, 2021) found that an optimism-based method did best overall, and that a purely greedy policy came a surprisingly close second when contexts were diverse enough to provide exploration for free. That result comes from supervised datasets converted to bandit problems, so it is suggestive only.

### Relation to A/B Testing and Reinforcement Learning

|  | A/B test | Contextual bandit | Full reinforcement learning |
| --- | --- | --- | --- |
| Allocation | Fixed split for the test's duration | Adapts continuously | Adapts continuously |
| Uses context | No (one winner overall) | Yes (a winner per context) | Yes |
| Models long-term effects of actions | No | No; each round is treated as independent | Yes; states and delayed rewards |
| Main output | A statistically clean estimate of an effect | A policy that earns reward while learning | A policy for sequences of decisions |

An A/B test is designed for inference: whether a change works and by how much. A bandit is designed for optimisation: losing as little as possible while finding out. Because a bandit shifts traffic toward winners, its data is not uniformly randomised and naive averages from it are biased. The comparison is developed in [Contextual Bandits vs A/B Testing](https://metavert.io/compare/contextual-bandits-vs-a-b-testing).

### Off-Policy Evaluation

A valuable property of bandits is that a new policy can be assessed on logs collected by an old one, without exposing users to it, provided the old policy randomised and recorded the probability of each choice. Li et al. (WSDM 2011) introduced a replay method that is provably unbiased when logged actions were chosen uniformly at random. Inverse propensity scoring generalises it to non-uniform logging by reweighting each record by one over its logged probability, at the cost of high variance. The doubly robust estimator (Dudík, Langford and Li, ICML 2011) combines a reward model with propensity weights and is accurate if either one is good. The Open Bandit Dataset and Pipeline (Saito et al., NeurIPS 2021) made real logged bandit data and standard estimators public. The practical corollary is to log, for every decision, the context, action, reward and action probability; the Vowpal Wabbit library encodes exactly that as `action:cost:probability | features`.

### Typical Uses and Pitfalls

Common applications are content [recommendation](https://metavert.io/recommendation-systems), [personalisation](https://metavert.io/personalization) of layouts and messages, choice among promotional offers, and creative selection in advertising.

- **Delayed or proxy rewards.** Optimising clicks is easy; optimising retention or revenue weeks later is not. A bandit trained on a short-term proxy will maximise the proxy.
- **Missing propensities.** Without logged action probabilities, off-policy evaluation is impossible and the logs are confounded.
- **No exploration floor.** If some action's probability falls to zero for a segment, the system can never discover that conditions changed.
- **Interference and carry-over.** If today's offer changes next week's behaviour, the independence assumption fails and the problem is really reinforcement learning.
- **Delegating exploration to a language model.** There is early evidence that LLM agents explore poorly when left to do it implicitly: Choi et al. (July 2026) report "myopic and polarized interaction patterns" when agents must probe one another's capabilities. That is one study in a multi-agent setting, but it supports the cautious design in which a language model proposes candidate actions and an explicit bandit allocates traffic among them.

## Related Topics

- [Contextual Bandits vs A/B Testing](https://metavert.io/compare/contextual-bandits-vs-a-b-testing) — optimisation versus inference
- [Reinforcement Learning](https://metavert.io/reinforcement-learning) — the general framework of which bandits are the one-step case
- [Uplift Modeling](https://metavert.io/uplift-modeling) — the offline, causal route to the same who-gets-what question
- [Recommendation Systems](https://metavert.io/recommendation-systems) — the original large-scale application
- [Personalization](https://metavert.io/personalization) — what bandits are most often used to deliver
- [Decision Models](https://metavert.io/decision-models) — models that choose actions rather than only predict
- [Contextual Bandits for Gaming](https://metavert.io/industry/contextual-bandits-for-gaming) — offers, difficulty and content selection in games
- [Contextual Bandits for Advertising & Marketing](https://metavert.io/industry/contextual-bandits-for-advertising-marketing) — creative and message selection
- [Churn Prediction](https://metavert.io/churn-prediction) — a prediction that only pays off once a decision policy acts on it

## Further Reading

- [A Contextual-Bandit Approach to Personalized News Article Recommendation](https://arxiv.org/abs/1003.0146) — Li, Chu, Langford, Schapire, WWW 2010
- [An Empirical Evaluation of Thompson Sampling](https://papers.nips.cc/paper_files/paper/2011/hash/e53a0a2978c28872a4505bdb51db06dc-Abstract.html) — Chapelle and Li, NeurIPS 2011
- [Thompson Sampling for Contextual Bandits with Linear Payoffs](https://arxiv.org/abs/1209.3352) — Agrawal and Goyal, arXiv, September 2012
- [Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms](https://arxiv.org/abs/1003.5956) — Li, Chu, Langford, Wang, WSDM 2011
- [Doubly Robust Policy Evaluation and Learning](https://arxiv.org/abs/1103.4601) — Dudík, Langford, Li, ICML 2011
- [A Contextual Bandit Bake-off](https://arxiv.org/abs/1802.04064) — Bietti, Agarwal, Langford, JMLR 2021
- [Open Bandit Dataset and Pipeline: Towards Realistic and Reproducible Off-Policy Evaluation](https://arxiv.org/abs/2008.07146) — Saito et al., NeurIPS 2021
- [Contextual Bandits and Vowpal Wabbit (tutorial)](https://vowpalwabbit.org/docs/vowpal_wabbit/python/latest/tutorials/python_Contextual_bandits_and_Vowpal_Wabbit.html) — Vowpal Wabbit documentation, accessed October 2026
- [Multi-Agent LLMs Fail to Explore Each Other](https://arxiv.org/abs/2607.11250) — Choi et al., arXiv, July 2026
