Uplift Modeling

Uplift modeling is the estimation, for each individual, of how much a treatment (a discount, a message, a retention call, a feature) changes the probability or size of an outcome compared with not treating them. It answers "whose behaviour will this action change?", where ordinary predictive models answer "who is likely to do it?". The quantity estimated is the conditional average treatment effect (CATE), and the field is also called incremental, net-lift or heterogeneous-treatment-effect modeling. Uplift models need data in which treatment was assigned at random, they cannot be scored against individual ground truth, and they are noisier than response models, but they are the correct tool whenever an intervention has a cost and its benefit varies across people.

Why It Differs from Response and Churn Prediction

A response model trained on treated customers ranks people by how likely they are to buy. Many of those would have bought anyway, so targeting them spends budget on no incremental effect. A churn model ranks customers by risk of leaving, yet the riskiest may be beyond persuasion. The standard illustration divides a population into four groups: persuadables, who act only if treated; sure things, who act either way; lost causes, who act in neither case; and sleeping dogs, for whom the treatment backfires, as when a retention offer reminds a subscriber to cancel. Only the first group repays treatment, and the last should be actively avoided. No model of the outcome alone can separate them.

The clearest empirical statement is Eva Ascarza's "Retention Futility" (Journal of Marketing Research, February 2018). Using field experiments at a wireless provider and a membership organisation, it argued that targeting customers by their sensitivity to the intervention reduces churn more effectively than targeting those at highest risk; customers at high risk and customers who respond are often different people.

Methods

The fundamental difficulty is that an individual's uplift is never observed: each person is either treated or not. Methods work around this in three ways, the grouping used in Gutierrez and Gérardy's 2017 review.

MethodIdeaNotes
S-learnerOne model of the outcome with treatment as an input feature; uplift is the difference between predictions with the flag on and offSimple; can shrink the effect toward zero if the model largely ignores the flag
T-learner (two-model)Separate outcome models for treated and control groups; uplift is their differenceEach model optimises its own outcome, so errors in the difference can be large
X-learnerExtends the T-learner by imputing individual effects, modelling them directly, and blending with propensity weights (Künzel et al., 2017)Designed for unbalanced groups, such as a small control holdout
Class transformationRecode the target so that a single classifier's output is a function of upliftConvenient with balanced randomisation and binary outcomes
Uplift trees and causal forestsTrees whose splits maximise the divergence between treated and control outcome distributions (KL, Euclidean, chi-square criteria)Model the effect directly; causal forests (Wager and Athey) add confidence intervals

The S-, T- and X-learners are meta-learners: wrappers around any base model, commonly gradient-boosted trees. R-learners and doubly robust learners add corrections that matter most when treatment was not randomly assigned. Open-source implementations include CausalML, maintained by Uber, and scikit-uplift.

Evaluation: Qini and AUUC

Because true individual effects are unobservable, uplift models are judged at group level. Customers in a randomised test set are ranked by predicted uplift; for each targeting depth (the top 10%, 20% and so on) the outcome rates of treated and control customers within that slice are compared, and the cumulative incremental outcome is plotted against depth. A good model's curve rises steeply and then flattens or falls as it reaches people with no or negative effect; a random ranking gives a straight diagonal.

The Qini coefficient, introduced by Nicholas Radcliffe and analysed formally by Surry and Radcliffe (2011), summarises that curve as the area between it and the diagonal, by analogy with the Gini coefficient for ordinary classifiers. AUUC, the area under the uplift curve, is a closely related summary; libraries differ in how they normalise it, so scores from different tools are not comparable. Both are noisy, because they rest on differences between subgroup rates, so bootstrap confidence intervals are worth reporting.

Data Requirements

Uplift modeling begins with an experiment. The cleanest training data is a randomised controlled trial in which a random subset of eligible customers receives the treatment and a random holdout does not, with the same features recorded for both before assignment. Several consequences follow.

  • Keep a permanent randomised holdout. Once a model decides who is treated, the resulting data is no longer random. A slice of traffic assigned by coin flip preserves the ability to retrain and to measure.
  • Expect to need more data than for prediction. The signal is a difference between two rates, often a few percentage points, so sample sizes that suffice for a response model can leave an uplift model fitting noise.
  • Use only pre-treatment features. Anything measured after assignment can be affected by the treatment and will bias the estimate.
  • Treat observational data with suspicion. Methods exist for estimating effects from non-random campaign history, but they rely on the assumption that every factor driving past targeting was recorded, which is seldom checkable.
  • Define the outcome in value terms. A discount that lifts conversion but lowers margin may have negative net uplift; the outcome should be net of the cost of the treatment where possible.

Uplift modeling is the offline counterpart of a contextual bandit: both decide who gets which action, one from a completed experiment and the other by experimenting continuously.

Further Reading