Behavioral Foundation Models

Behavioral foundation models are large sequence models, usually transformers, pretrained with self-supervision on logs of user events (purchases, payments, sessions, clicks, in-app actions) so that a single model can be reused for many downstream predictions about those users. They apply the foundation-model pattern to behaviour rather than text: learn general representations from raw sequences once, then adapt them to churn, lifetime value, fraud, credit or recommendation tasks. As of October 2026 the public evidence is real but narrow. Most of it comes from a few financial-services companies describing their own systems, and published head-to-head comparisons against well-tuned gradient-boosted trees are rare.

The term needs one disambiguation. In reinforcement-learning research, "behavioral foundation model" is also used for agents pretrained to produce many behaviours in an environment. This page concerns the other sense: models of what people do, learned from event streams.

How They Are Built

An event is not a word. It is a small record with a timestamp, a type, an amount, a merchant or item, perhaps free text. The first design problem is therefore tokenisation: turning each record into something a transformer can consume. Two approaches appear in the literature. One serialises the fields into tokens, as language models do. The other uses an event embedder, a small network that maps each multi-field event to a single vector, followed by a sequence model over those vectors.

The pretraining objective is self-supervised. Decoder-style models predict the next event; encoder-style models reconstruct masked events. The resulting user representation is then used in one of three ways: as embeddings fed to a simple downstream model, as extra columns alongside existing engineered features in a tabular model, or by fine-tuning the whole network for one task.

Published Examples

WorkDesignWhat the authors report
nuFormer (Braithwaite et al., Nubank, July 2025)Transformer over transaction sequences mixing text and structured fields; embeddings fine-tuned jointly with tabular featuresGains on a large-scale recommendation task, attributed to better representations rather than new data sources
PRAGMA (Ostroukhov et al., Revolut, April 2026)Encoder family pretrained with masked modelling on banking event historiesStrong results from a simple linear model on the embeddings, across tasks including fraud detection, credit assessment and lifetime-value estimation
Rusakov et al. (July 2026)Early fusion of transactions and digital-interaction events; next-event pretrainingRepresentations combined with existing features outperform task-specific models; deployed at a large Eastern European bank with "measurable improvements in business metrics"

All three are self-reported by the organisations that built the systems, on private data that cannot be reproduced, and the abstracts give few numbers that would let an outsider compare them with one another or with a strong tree baseline. A pattern does recur, though: the learned representation is added to engineered features rather than replacing them.

Scaling Laws

The first systematic study of how to spend compute on such models is "Scaling Laws for Behavioral Foundation Models over User Event Sequences" (Brüel Gabrielsson, June 2026). It reports about 600 training runs spanning 1015 to 1019 training FLOPs on recommendation, payments and commerce data, using a feature-based event embedder and a decoder-only transformer trained on next-event prediction. Its findings differ from the text-model playbook in several ways. The compute-optimal embedder is small, around 2% of total parameters. Training is data-heavy relative to text at low compute, with the ratio moving towards language-model norms as budgets grow. The best number of sampled negatives rises with compute until memory, not arithmetic, becomes the constraint. And the choice of evaluation metric changes the optimal recipe; in the author's words, "the evaluation metric is therefore part of the scaling law."

This is a single paper by a single author. It gives practitioners a starting recipe, not a settled law, and it says nothing about whether a model trained in one domain transfers to another.

Uses

The practical appeal is consolidation. A company that maintains dozens of hand-built feature pipelines, one per model, can in principle replace much of that work with one representation of each user that is refreshed as events arrive. The applications named in the published work are fraud and credit risk, product recommendation, engagement and retention prediction, and lifetime-value estimation. The same idea carries over to any product with dense event logs, including subscription media, commerce and live-service games, although public evidence outside financial services is thin.

Open Questions

Several questions are unresolved. Do they beat trees? Published results compare against the authors' internal baselines; none of the work cited here reports a controlled comparison with tuned gradient-boosted trees or tabular foundation models on shared data. Do they transfer? Unlike text, event schemas are specific to each company, so today's models are foundation models for one organisation's data, not for behaviour in general. What do they cost? Pretraining and serving a sequence model for every user is far more expensive than scoring a tree. How are they governed? Learned embeddings of personal financial or behavioural histories raise privacy, explainability and fairness questions that hand-built features, for all their cost, make easier to audit. Finally, a better prediction is not a better decision: turning a churn or value score into an action still needs uplift modeling or a contextual bandit.

Further Reading