Tabular Foundation Models

Tabular foundation models are neural networks pretrained once, across millions of tables, so that they can make predictions on a new spreadsheet-style dataset without being trained on it: the labelled rows are passed in as context and the model outputs predictions for the unlabelled rows in a single forward pass. They are the first serious challenge in two decades to gradient-boosted decision trees such as XGBoost, LightGBM and CatBoost on structured data, and as of October 2026 they lead the main public benchmark for small and medium tables. They have not made boosted trees obsolete: the strongest models carry licence restrictions, need a GPU, and move the computational cost from training to prediction time.

How They Work: Prior-Data Fitted Networks

The dominant recipe is the prior-data fitted network (PFN). Instead of collecting real tables, the developers define a prior, a generator of synthetic prediction problems, and train a transformer on a very large number of them to predict held-out labels given the rest of the table. The original TabPFN paper (Hollmann, Müller, Eggensperger and Hutter, July 2022) describes the result as a model that takes training and test samples as one set-valued input and returns predictions in a single forward pass. This is in-context learning, the same mechanism by which large language models use examples in a prompt, applied to rows and columns. There is no gradient descent on the user's data and, by default, no hyperparameter search.

The Nature paper on TabPFN v2 (Hollmann et al., January 2025) reports pretraining on roughly 130 million synthetic datasets generated from structural causal models, with an attention scheme in which each cell attends first along its row and then along its column. Not every model in the class is purely synthetic: TabDPT (Ma et al., NeurIPS 2025) pretrains on real tables and argues that real data carries signal synthetic priors miss.

The Main Models

ModelOriginNotable property (as reported by its authors)
TabPFN familyUniversity of Freiburg, then Prior Labsv2 (Nature, 2025) targeted up to 10,000 rows; TabPFN-3.5 (September 2026) accepts up to 1,000,000 rows and 20,000 features
TabICL / TabICLv2Qu, Holzmüller, Varoquaux, Le MorvanColumn-then-row attention compresses each row to a fixed embedding before in-context learning; v2 (ICML 2026) releases inference and pretraining code
Mitra-v2Amazon researchers77-million-parameter model trained only on synthetic data; weights and code under Apache-2.0
TabDPTMa et al.Retrieval-based in-context learning pretrained on real tables; open weights and training code

Benchmark Standing

The reference benchmark is TabArena (Erickson et al., NeurIPS 2025 Datasets and Benchmarks), a continuously maintained leaderboard built from 51 curated datasets with between 500 and 250,000 training rows. Its paper concluded that "tabular foundation models dominate for small data," that gradient-boosted trees "are still strong contenders," and that tuned, ensembled deep-learning models had caught up with them. It also recorded an awkward detail: the foundation models of the time could not run everywhere, so TabPFN v2 was scored on 33 of the datasets and TabICL on 36.

Since then the leaderboard has changed hands several times, and almost every claim comes from the model's own authors. TabICLv2 (February 2026) reported surpassing the then-leading RealTabPFN-2.5 without tuning. Prior Labs' TabPFN-3 report (May 2026) said a single forward pass beat all other models, including tuned and ensembled baselines. Its TabPFN-3.5 release (September 2026) states that the model "ranks first place on both TabArena and BeyondArena," the latter a 142-dataset suite the company assembled around grouped, temporal, high-cardinality and text-rich data. These are vendor-reported results; the leaderboard is public, but readers should check the live standings rather than rely on any release note.

Limits

Three constraints matter in practice. The first is size. Row and feature ceilings have risen quickly, from 1,000 rows in 2022 to a million in 2026, but in-context learning still holds the whole training set in memory at prediction time; the Nature paper notes that memory grows with dataset size. The second is latency and hardware. A boosted tree is slow to tune and fast to score; a tabular foundation model reverses that. The TabPFN repository recommends a GPU and limits CPU use to about 5,000 samples for its current models. Distilling the model into a small network or tree ensemble is the usual answer for real-time serving. The third is licensing. The code is often open while the best weights are not: TabPFN-2.5 and later are released under non-commercial licences, whereas TabPFN v2, TabICLv2 and Mitra-v2 are more permissive. Being open-weight is not the same as being free to deploy.

When Gradient-Boosted Trees Still Win

Boosted trees remain the sensible default when the table has many millions of rows, when predictions must be served on CPUs within milliseconds, when a permissive licence is mandatory, or when a tuned and monitored model already exists and the expected gain is small. They also come with mature tooling for monotonic constraints, explanations and incremental retraining. Tabular foundation models are strongest where data is scarce, where tuning time is the bottleneck, or where many small models must be built quickly, which is why they are being tested for tasks such as churn prediction and lifetime value estimation. The honest procedure is a head-to-head comparison on one's own data with time-ordered splits; public leaderboards measure average rank across datasets, not performance on a particular table. The trade-offs are set out in TabPFN vs XGBoost.

Further Reading