Tabular Foundation Models for Retail

Industry Application
Tabular Foundation ModelsRetail

Tabular foundation models for retail are pretrained in-context models such as TabPFN used on the structured data retailers already hold (sales histories by product and store, customer and basket tables, price and promotion logs) to forecast demand, score purchase propensity and inform pricing without fitting a separate model for each task. Retail is a mixed fit. Many of its tables are far larger than these models were designed for, while some of its hardest problems, such as new products, new stores and customers with short histories, are small-data problems in disguise. Independent retail-specific evidence is limited as of October 2026, and most of what exists is academic or vendor-reported.

Demand forecasting

Retail demand forecasting has a well-known reference point in the M5 accuracy competition (Makridakis, Spiliotis and Assimakopoulos, International Journal of Forecasting, 2022), which asked participants to forecast 42,840 hierarchical unit-sales series from Walmart. Any new method is implicitly measured against that tradition of feature-engineered, per-retailer models.

The bridge from tables to forecasting is TabPFN-TS (Hoo, Müller, Salinas and Hutter, first posted January 2025, revised January 2026). It recasts forecasting as tabular regression by adding lightweight temporal features to a pretrained TabPFN-v2, needs no time-series pretraining, supports covariates such as price or promotion flags, and with about 11 million parameters is reported by its authors as competitive on the GIFT-Eval and fev-bench benchmarks. Those are general forecasting benchmarks, not retail ones. Prior Labs states that TabPFN-3.5 (September 2026) targets grouped, multi-location and temporal data, a description that matches store-by-product panels, but this is a vendor positioning statement and no public M5-style comparison accompanies it.

The clearest retail use is where history is short. A product launched last month or a store opened last quarter has too few observations to train a dedicated model, and a pretrained model conditioned on comparable items is a reasonable candidate. For a long-running assortment at national scale, a tuned gradient-boosted pipeline has the advantages of scale, cost and years of accumulated feature engineering.

Propensity, churn and choice

For customer-level tables the best cross-industry evidence is the churn benchmark presented at the ICML 2026 workshop on foundation models for structured data. Across nine public datasets from seven sectors, TabICL v2 ranked in the top three on eight and beat XGBoost by up to 9.23 percentage points of PR-AUC with no dataset-specific retraining. The abstract does not list the sectors, so the share that is retail is unknown.

A more retail-specific result comes from marketing science. Liu and Zhang (arXiv, July 2026) asked whether tabular foundation models can estimate discrete choice, the standard demand framework in which a shopper picks one brand from a set. They identify a structural mismatch: the models assume independent rows, whereas choice depends on the whole set on offer and on persistent differences between consumers. Their reformulation encodes both, and on a yogurt scanner panel it exceeded hierarchical Bayesian estimation in predictive accuracy while running 16 times faster, with the largest gains for consumers with 10 to 40 recorded purchases. It is one dataset in one category, and the result depends on the reformulation, not on feeding the raw panel to the model.

Product descriptions and reviews are a further angle. Prior Labs reports that TabPFN-3.5-Plus, the variant for text-rich columns, gains up to about 250 Elo over the strongest previous baseline on text-rich and high-cardinality data in its own BeyondArena evaluation. High-cardinality categorical columns such as SKU, brand and supplier are typical of retail tables, which makes the claim relevant, but it remains vendor-reported.

Pricing and promotions

Pricing is where prediction and decision most need separating. A model that predicts units sold at observed prices learns from prices the retailer chose, usually in response to expected demand, so its implied price sensitivity is confounded. Estimating what would happen at a different price is a causal question better handled with experiments, uplift modelling or contextual bandits, with the tabular model supplying forecasts or features. No public study was found that validates a tabular foundation model for price-elasticity estimation.

Regulation is moving on the customer-specific end of pricing. On 19 August 2026 the US Federal Trade Commission sought comment on a proposed enforcement policy statement on personalised pricing, which it defines as using personal data to set prices according to what a company believes an individual is willing to spend, and warned that undisclosed use of personal data for that purpose could violate the FTC Act. Comments closed on 18 September 2026 and the statement had not been finalised when this page was written. Propensity scores built on customer tables are exactly the kind of input such a policy would reach if they feed individual prices.

Scale and cost

The TabPFN-3 report (May 2026) gives a vendor-stated envelope of 1 million training rows and 200 features on one H100 GPU. A daily store-by-product table for a mid-sized chain passes that within weeks, so practical use means sampling, aggregating to a higher level of the hierarchy or restricting the model to segments with little data. The repository also lists the recent weights under non-commercial licences as of October 2026, which makes a commercial agreement or the hosted API a precondition for production use.

Applications & Use Cases

New-product and new-store forecasts

Short-history items conditioned on comparable products or locations, where there is too little data to train a dedicated model.

Covariate-aware demand baselines

TabPFN-TS accepts price and promotion covariates, giving planners a quick baseline to set against the production forecaster.

Purchase and churn propensity

Customer-table scoring for retention and campaign targeting; cross-sector churn benchmarks favour in-context models, with retail-specific shares unreported.

Brand choice with short purchase histories

Reformulated choice estimation performed best for shoppers with 10 to 40 purchases in the one published scanner-panel study.

Text-rich catalogue tables

Models with native handling of descriptions and reviews avoid a separate embedding pipeline; evidence is vendor-reported.

Supplier and returns risk

Smaller operational tables, such as supplier delay or return likelihood, sit comfortably within the size limits.

Key Players

  • Prior Labs — Developer of TabPFN, TabPFN-TS and the text-oriented TabPFN-3.5-Plus; lists churn, lifetime value, segmentation and pricing among its use cases.
  • TabICL — In-context tabular model from Qu et al. (ICML 2025); its v2 led the 2026 churn benchmark.
  • Walmart — Source of the M5 competition sales data, the standard public retail forecasting dataset.
  • M Competitions — Forecasting competitions organised by Spyros Makridakis and colleagues; M5 results were published in 2022.
  • US Federal Trade Commission — Proposed an enforcement policy statement on personalised pricing in August 2026.
  • LightGBM and XGBoost — Gradient-boosted tree libraries that remain the production default for large-scale retail forecasting and propensity models.

Challenges & Considerations

  • Table size — Item-store-day panels exceed the stated 1 million-row envelope quickly; sampling and aggregation choices then drive results as much as the model does.
  • Hierarchy and coherence — Retail forecasts must add up across product and location hierarchies. Row-wise models do not guarantee this, so reconciliation is still needed.
  • Prediction is not elasticity — Observed prices are confounded with expected demand. Using a forecasting model to set prices without experimental variation risks systematic error.
  • Personalised-pricing scrutiny — The FTC's proposed statement targets undisclosed use of personal data to set individual prices; propensity models that feed pricing fall within its subject matter.
  • Thin independent evidence — Retail results come from one scanner-panel study, general benchmarks and vendor reports. Licensing terms for current weights add a commercial hurdle.