# Contextual Bandits vs A/B Testing

> Contextual bandits adapt traffic per user to maximise reward; A/B tests use fixed random splits to measure causal effect. How to choose between them.

Source: https://metavert.io/compare/contextual-bandits-vs-a-b-testing  
Published: 2026-10-07  
Updated: 2026-10-07

Comparison

**Contextual bandits vs A/B testing** is a choice between two ways of spending traffic: an A/B test splits users at random between fixed variants to measure which is better on average, while a [contextual bandit](https://metavert.io/contextual-bandits) learns a policy that picks a variant for each user from their features and shifts traffic toward what is working as data arrives. The short answer is that they answer different questions. An A/B test is a measurement instrument; a contextual bandit is a decision system that happens to produce data.

That difference decides most real cases. If the goal is a defensible estimate of whether a change helped, randomised assignment with a fixed design is still the cleanest tool. If the goal is to keep choosing well among many options whose value differs by user and decays over time, a bandit wastes less traffic on losers. Many mature teams run both: bandits for allocation, with a randomised holdout kept aside so the whole system can still be measured.

## Feature Comparison

| Dimension | Contextual Bandits | A/B Testing |
| --- | --- | --- |
| Primary goal | Maximise cumulative reward while learning (minimise regret) | Estimate the causal effect of a change with known error rates |
| Assignment | Adaptive: probability of each action depends on context and past outcomes | Fixed random split, independent of user features and of results so far |
| Unit of decision | Per user or per request, conditioned on a feature vector | Per population: one winner is shipped to everyone (or to pre-defined segments) |
| Output | A policy mapping context to action, updated continuously | A point estimate and confidence interval for each metric |
| Personalisation | Native: different users can receive different actions | Only through pre-planned segment analysis, with multiple-comparison risk |
| Statistical inference | Harder: adaptively collected data bias naive averages (Nie et al., 2017); needs logged propensities and off-policy estimators | Standard: classical tests apply if the sample size is fixed in advance; repeated peeking invalidates them (Johari et al., 2015) |
| Number of variants | Scales to many actions; poor arms are starved of traffic early | Each extra arm divides traffic and lengthens the test |
| Non-stationarity | Can keep exploring and adapt as preferences drift | Result is a snapshot of the test period; re-test to detect drift |
| Reward signal needed | Fast, attributable, per-decision feedback (clicks, conversions within minutes or hours) | Works with slow or long-horizon metrics if the test simply runs longer |
| Engineering cost | Feature pipeline, propensity logging, online or frequent retraining, policy deployment loop | Randomisation service, metric pipeline, analysis tooling |
| Typical failure mode | Incorrect logging, feedback loops, optimising a proxy reward | Underpowered tests, peeking, novelty effects, shipping an average winner that loses for a segment |
| Auditability | Decisions vary by user and time; explanation requires the logged context and policy version | One documented decision per experiment |

## Detailed Analysis

### Two different objectives

The controlled experiment is built to isolate cause. Random assignment makes the treatment and control groups comparable, so a difference in the chosen metric can be attributed to the change. The practical literature that grew out of large web experimentation programmes, notably Kohavi, Henne and Sommerfield's survey and practical guide, adds the discipline around it: a single agreed evaluation criterion, sample sizes fixed from a power calculation, and A/A tests to check the instrument itself.

A contextual bandit optimises something else: the reward earned during the experiment. Li, Chu, Langford and Schapire's 2010 paper, which introduced LinUCB for news recommendation, framed the problem as sequentially choosing articles from user and article features while adapting to click feedback. On a Yahoo! Front Page Today Module dataset of more than 33 million events they reported a 12.5% click lift over a context-free bandit, with a larger advantage when data was scarce. The comparison in that paper is against a non-contextual bandit rather than against an A/B test, which is worth remembering when the figure is quoted.

### What adaptivity costs in inference

Shifting traffic toward apparent winners is exactly what breaks textbook statistics. Nie, Tian, Taylor and Zou (2017) showed that sample means from adaptively collected data are systematically biased downward, because arms that look bad early receive few further samples and never get the chance to regress upward. A bandit's dashboard is therefore not a trustworthy estimate of each arm's true effect.

A/B tests have their own version of the problem. Johari, Pekelis and Walsh (2015) pointed out that standard analyses are unreliable when experimenters continuously monitor results and stop when they like what they see, and proposed always-valid p-values and confidence intervals as a remedy. Sequential testing narrows the gap between the two approaches: it permits early stopping without giving up error control, but it still allocates traffic evenly and still returns one answer for the whole population.

### Evaluation without a live test

Bandit practice depends on logging the probability with which each action was chosen. With those propensities, a new policy can be evaluated offline against historical data. Li et al. proposed an offline evaluation method using randomly served traffic; Dudík, Langford and Li (2011) introduced doubly robust estimation, which stays accurate when either the reward model or the model of the logging policy is good, and reported lower variance than earlier estimators. This is a real advantage over A/B testing, where every new idea costs a fresh slice of live traffic. It only works if exploration never drops to zero and the logs are correct.

### Evidence from production systems

Agarwal and co-authors' description of Microsoft's Decision Service (2016) is the most detailed public account of running contextual bandits as a product. It organises the system as a loop of explore, log, learn and deploy, and names incorrect data collection and weak debuggability as the failures that sink most attempts. The authors report click-through improvements of 25–30% in content recommendation and an 18% revenue lift in landing-page optimisation. These are author-reported results from their own deployments, not independent replications.

The research record is also less one-sided than vendor material suggests. In Bietti, Agarwal and Langford's "Contextual Bandit Bake-off" (JMLR), run across a large set of supervised-learning datasets, an optimism-based method performed best overall, but a simple greedy baseline that explores only through the natural diversity of contexts was a surprisingly strong competitor. Sophisticated exploration does not always pay for itself.

### Using them together

The common production pattern layers the two. A bandit allocates traffic within an experience, and an outer A/B test or a persistent random holdout compares "bandit on" against a fixed baseline, which restores a clean causal estimate of the system's total value. Where the question is who benefits from an intervention rather than which variant is best, [uplift modelling](https://metavert.io/uplift-modeling) on randomised data is often a better fit than either. One caution for teams building [agent systems](https://metavert.io/multi-agent-systems): Choi et al. (July 2026) found that LLM agents explore poorly when choosing among peers, showing myopic and polarised interaction patterns that raise regret. It is a single paper in a multi-agent setting, but it supports keeping exploration in an explicit algorithm rather than leaving it to a language model's judgement.

## Best For

#### Launch decision for a redesign or pricing change

A/B Testing

A one-off, high-stakes decision needs an unbiased effect size and an interval that stakeholders can audit. Fixed randomisation gives that directly.

#### Choosing among dozens of creatives or offers

Contextual Bandits

With many arms, an even split spends most traffic on poor options. A bandit starves weak arms early and concentrates on contenders.

#### Per-user content or offer selection

Contextual Bandits

When the best action differs by user, a single global winner leaves value unclaimed. This is the setting Li et al. studied for news recommendation.

#### Metrics that mature over weeks (retention, lifetime value)

A/B Testing

Bandits need quick, attributable rewards. Slow outcomes are better measured with a fixed test, or with a validated short-term proxy.

#### Low-traffic products

Depends

Neither method creates data. A/B tests become underpowered; contextual policies with many features overfit. Fewer variants and a simpler, non-contextual bandit may be the honest option.

#### Short-lived campaigns and seasonal events

Contextual Bandits

If the campaign ends before a test reaches significance, learning while serving is the only way to benefit from the data collected.

#### Regulated or contested decisions

A/B Testing

A pre-registered test with one documented outcome is easier to defend than a policy whose decisions vary by user and by hour.

#### Always-on optimisation with accountability

Both

Run the bandit for allocation and keep a randomised holdout so the uplift of the whole system stays measurable.

## The Bottom Line

Contextual bandits and A/B tests are complements more often than rivals. An A/B test is the right tool when the deliverable is knowledge: a causal estimate with stated error rates that will justify a decision. A contextual bandit is the right tool when the deliverable is a stream of good decisions across many options, for users who differ, in an environment that drifts.

The costs are asymmetric. A/B testing is cheap to run and expensive in traffic; bandits are economical with traffic and expensive in engineering, because correct propensity logging, reward attribution and off-policy evaluation are prerequisites rather than refinements. The published evidence also argues for modest expectations: the best-known lift figures are author-reported, and simple greedy baselines hold up well in benchmark comparisons.

A reasonable default is to start with disciplined A/B testing, add sequential methods when early stopping matters, and adopt contextual bandits where there are many actions, fast rewards and real heterogeneity across users, keeping a random holdout so the bandit itself remains testable.

## Related Topics

- [Contextual Bandits](https://metavert.io/contextual-bandits)
- [Uplift Modeling](https://metavert.io/uplift-modeling)
- [Reinforcement Learning](https://metavert.io/reinforcement-learning)
- [Recommendation Systems](https://metavert.io/recommendation-systems)
- [Personalization](https://metavert.io/personalization)
- [Decision Models](https://metavert.io/decision-models)
- [Contextual Bandits for Gaming](https://metavert.io/industry/contextual-bandits-for-gaming)
- [Contextual Bandits for Advertising & Marketing](https://metavert.io/industry/contextual-bandits-for-advertising-marketing)

## Further Reading

- [A Contextual-Bandit Approach to Personalized News Article Recommendation (Li, Chu, Langford, Schapire, 2010) – arXiv](https://arxiv.org/abs/1003.0146)
- [Controlled Experiments on the Web: Survey and Practical Guide (Kohavi, Henne, Sommerfield) – ExP Platform](https://exp-platform.com/Documents/GuideControlledExperiments.pdf)
- [Always Valid Inference: Bringing Sequential Analysis to A/B Testing (Johari, Pekelis, Walsh, 2015) – arXiv](https://arxiv.org/abs/1512.04922)
- [Why Adaptively Collected Data Have Negative Bias and How to Correct for It (Nie, Tian, Taylor, Zou, 2017) – arXiv](https://arxiv.org/abs/1708.01977)
- [Doubly Robust Policy Evaluation and Learning (Dudík, Langford, Li, 2011) – arXiv](https://arxiv.org/abs/1103.4601)
- [Making Contextual Decisions with Low Technical Debt (Agarwal et al., 2016) – arXiv](https://arxiv.org/abs/1606.03966)
- [A Contextual Bandit Bake-off (Bietti, Agarwal, Langford) – arXiv / JMLR](https://arxiv.org/abs/1802.04064)
- [Multi-Agent LLMs Fail to Explore Each Other (Choi et al., July 2026) – arXiv](https://arxiv.org/abs/2607.11250)
