Contextual Bandits for Advertising & Marketing
Contextual bandits for advertising and marketing are online learning policies that pick a creative, message, offer or budget split for each impression, recipient or campaign using the available context, observe the response, and shift future choices toward what works while continuing to explore. Advertising is where the method was first proven at scale and where it has the longest published record. That record has a clear shape: detailed papers from large platforms with their own traffic, and much less from ordinary advertisers, whose bandits usually run inside ad platforms they cannot inspect.
The published record
The founding result is Li, Chu, Langford and Schapire (2010), who modelled article selection on the Yahoo front page as a contextual bandit, introduced the LinUCB algorithm and reported a 12.5% click lift over a context-free bandit on more than 33 million events. Microsoft followed with the Decision Service (Agarwal et al., 2016), described as the first general system for contextual learning, and reported click-through improvements of 25 to 30% in two content-recommendation deployments and an 18% revenue lift on a landing page. Spotify's "Explore, Exploit, Explain" (McInerney et al., RecSys 2018) used a bandit to choose jointly which content to recommend and which explanation to attach, reporting a significant engagement improvement in live tests without a headline number in the abstract.
In advertising proper, Du et al. (KDD 2021) combined Gaussian-process uncertainty estimates with deep click models to drive exploration, and reported an 8.2% gain in social welfare and 8.0% in revenue in an online A/B test on Alibaba's display advertising platform. All of these figures are reported by the companies that built the systems, on their own traffic and metrics. They show that the method works at platform scale; they are not estimates of what a typical advertiser should expect.
Creative and offer selection
The most recent detailed account is from Uber (March 2025), which applied contextual bandits to CRM email. Each combination of content elements is an arm, and the system offers both LinUCB and an XGBoost reward model with SquareCB exploration, which assigns each action a probability that falls with its gap from the best predicted score. Subject lines and pre-headers are turned into features using 1,536-dimensional text embeddings reduced to 128 dimensions by PCA, which lets the model generalise to variants it has never sent. Uber states that experiments with 100 or more variants can converge in weeks where A/B tests of two or three variants took four to six, and publishes no lift figures. The embedding step is the notable design choice: it turns creative selection from a problem with one parameter per creative into one where new creatives start with an informed estimate.
Budget allocation
Budget is a harder action space because it is continuous, constrained and shared between campaigns. Amazon researchers (Ge et al., AdKDD workshop at KDD 2024) formulated allocation across campaigns and ad lines as a multi-task combinatorial bandit, using a Bayesian hierarchical model to share information between campaigns through their metadata and Thompson sampling to explore, and evaluated it offline and in online experiments on Amazon campaign data. The abstract reports no numeric results. Hierarchical sharing is the important idea for practitioners, because individual campaigns rarely generate enough conversions to learn alone.
Logging, replay and holdouts
Measurement is where bandits differ most from the A/B tests marketers already run. Because allocation changes during the test, naive comparisons between arms are biased. Li et al. showed in a companion paper that if some traffic is served uniformly at random, a new policy can be evaluated offline by replay with provably unbiased results, which they checked against online bucket tests at Yahoo. The Decision Service paper generalised the lesson: its explore and log abstractions exist to guarantee that every decision is recorded with its context and probability, since faulty logging was the common source of technical debt. For a marketing team the practical translation is a persistent random holdout and propensity logging from the first day, without which the lift attributed to the bandit cannot be separated from seasonality or audience drift.
Limits and constraints
Three constraints recur. Conversions are delayed and sparse compared with clicks, so policies either optimise a fast proxy or learn slowly. Tooling has thinned: Microsoft retired Azure AI Personalizer on 25 August 2026 and directs users to its open-source learning-loop project, which moves operational responsibility in-house. Its documentation had advised at least about 1,000 events per day and at most about 50 actions per decision, a reminder that the method needs volume. And context is personal data. Research on differentially private contextual bandits (Shariff and Sheffet, 2018) shows privacy can be built in at a quantified cost in regret, and the FTC's August 2026 proposal on personalised pricing signals that offer policies which vary price by individual, as opposed to message or creative, face disclosure expectations in the United States.
Applications & Use Cases
Email and push creative selection
Choosing subject line, pre-header and content block per recipient from a large variant pool, with text embeddings as action features, as in Uber's published system.
Display and feed creative rotation
Allocating impressions among creatives by context with uncertainty-driven exploration, the setting of the Alibaba display-advertising study.
Landing-page and on-site modules
Selecting page layouts or recommended modules; the Decision Service paper reports an 18% revenue lift on one landing page.
Cross-campaign budget allocation
Splitting a fixed budget over campaigns and ad lines with Thompson sampling and hierarchical sharing between related campaigns.
Offer and incentive choice
Picking which promotion a customer sees, combined with uplift modelling to avoid discounting customers who would have bought anyway.
Recommendation with explanations
Jointly choosing the item and the reason shown for it, as in Spotify's bandit for explained recommendations.
Key Players
- Yahoo — Site of the original LinUCB study and of the replay method for offline bandit evaluation.
- Microsoft — Built the Decision Service and Azure AI Personalizer (retired August 2026); Microsoft Research is a major contributor to Vowpal Wabbit.
- Spotify — Published the 2018 bandit approach to explained recommendations tested on production traffic.
- Alibaba — Its display advertising platform hosted the 2021 uncertainty-aware exploration study.
- Amazon — Published the 2024 multi-task combinatorial bandit for advertising budget allocation.
- Uber — Described in 2025 a contextual-bandit platform for personalised CRM communication.
- Vowpal Wabbit — Open-source online learning library used for contextual bandit and reinforcement learning implementations.
Challenges & Considerations
- Self-reported results — Every lift figure in the public record comes from the operator of the system. Selection toward successful projects is likely, and none is an independent replication.
- Delayed and sparse conversions — Purchases arrive days after exposure and are rare, which slows learning or pushes teams toward click proxies that may not track revenue.
- Opaque platform optimisation — Advertisers buying through large ad platforms generally cannot see action probabilities or exploration settings, which prevents their own off-policy evaluation.
- Privacy and context features — Contextual policies consume personal data. Privacy-preserving variants exist in research at a cost in performance, and consent rules limit which features are usable.
- Personalised pricing — Varying price by individual attracts regulatory attention that varying creative does not; the FTC's 2026 proposal treats undisclosed use of personal data for pricing as a potential FTC Act violation.
Further Reading
- A Contextual-Bandit Approach to Personalized News Article Recommendation (Li, Chu, Langford and Schapire) — arXiv, February 2010
- Unbiased Offline Evaluation of Contextual-bandit-based News Article Recommendation Algorithms (Li et al.) — arXiv, March 2010
- Making Contextual Decisions with Low Technical Debt (Agarwal et al.) — arXiv, June 2016 (revised May 2017)
- Explore, Exploit, Explain: Personalizing Explainable Recommendations with Bandits (McInerney et al.) — Spotify Research / RecSys 2018
- Exploration in Online Advertising Systems with Deep Uncertainty-Aware Learning (Du et al.) — arXiv / KDD 2021, November 2020
- Enhancing Personalized CRM Communication with Contextual Bandit Strategies — Uber Engineering Blog, March 2025
- Multi-task combinatorial bandits for budget allocation (Ge et al.) — Amazon Science / AdKDD 2024
- What is Personalizer? (archived documentation with retirement notice) — Microsoft Learn, read October 2026
- Differentially Private Contextual Linear Bandits (Shariff and Sheffet) — arXiv, September 2018
- FTC Seeks Comment on Enforcement Policy Statement Regarding Personalized Pricing — Federal Trade Commission, August 2026