TabPFN

TabPFN (Tabular Prior-data Fitted Network) is a family of pretrained transformer models for classification and regression on tables, developed in Frank Hutter's group at the University of Freiburg and now by the company Prior Labs. A TabPFN model is not fitted to a dataset in the usual sense: it reads the labelled rows as context and predicts the unlabelled ones in a single forward pass. It is the best-known tabular foundation model, and as of October 2026 the current release is TabPFN-3.5 (15 September 2026), which Prior Labs reports as first on the TabArena benchmark. Its code is Apache-2.0, but the weights of every version since TabPFN-2.5 are licensed for non-commercial use only.

How It Works

TabPFN is trained once, offline, on synthetic data. The developers specify a prior over data-generating processes, built on structural causal models, sample an enormous number of small artificial prediction problems from it, and train a transformer to predict masked labels in each one. The Nature paper on TabPFN v2 (Hollmann et al., January 2025) reports about 130 million synthetic datasets and roughly two weeks of training on eight consumer-grade GPUs, and summarises the result as "a learning algorithm that is itself learned." At inference the network approximates Bayesian prediction under that prior: the user's training rows and test rows go in together, and predictions come out without any gradient steps on the user's data.

The architecture treats a table as a grid. In v2 each cell attends to the other features in its row and then to the same feature across all rows, which makes the model indifferent to row and column order. The same paper notes that, as a generative model, TabPFN also supports fine-tuning, synthetic data generation, density estimation and reusable embeddings.

Versions

VersionDateStated scaleWeights licence
TabPFN (v1)July 2022 (arXiv)Up to 1,000 training samples, 100 numerical features, 10 classes; classification onlyOpen (released with the paper)
TabPFN v2January 2025 (Nature)Up to 10,000 samples and 500 features; adds regressionPrior Labs License (Apache 2.0 plus attribution)
TabPFN-2.5November 2025Up to 50,000 rows and 2,000 featuresNon-commercial
TabPFN-3May 2026Up to 1,000,000 training rows; benchmarked to 200 features at that sizeNon-commercial
TabPFN-3.5September 2026Up to 1,000,000 rows and 20,000 featuresNon-commercial

The headline claims are the developers' own. The v2 paper states that in 2.8 seconds TabPFN outperformed an ensemble of the strongest baselines tuned for four hours on classification datasets of up to 10,000 samples. TabPFN-2.5 reported a 100% win rate against default (untuned) XGBoost on small and medium classification datasets and parity with AutoGluon 1.4's four-hour ensemble. TabPFN-3 reported beating gradient-boosted trees tuned for eight hours on datasets of up to a million rows, running up to 20 times faster than 2.5, and introduced test-time compute scaling, a "Thinking" mode that spends more computation per prediction.

TabPFN-3.5 comes in four variants: the base model; Fast, an alpha that Prior Labs describes as up to six times faster than the base model; Plus, with added handling of text-heavy columns; and Thinking, the most accurate. The company reports the base model as first among more than 27 methods on TabArena's 51 datasets, about 150 Elo points ahead of the previous leader on BeyondArena (its own 142-dataset suite), and strongest on grouped, temporal, high-cardinality and text-rich tables. The release material contains no table comparing it with individually tuned XGBoost, LightGBM or CatBoost on named datasets, so the size of the advantage on any particular problem is something a user has to measure.

Using It

The open-source tabpfn Python package (Python 3.10 or later) follows the scikit-learn interface and loads TabPFN-3.5 by default:

from tabpfn import TabPFNClassifier

clf = TabPFNClassifier()
clf.fit(X_train, y_train)      # stores the context; no gradient training
pred = clf.predict(X_test)

A TabPFNRegressor is the counterpart for numeric targets, and earlier versions can be selected explicitly. The repository recommends a GPU (around 8 GB of memory is enough for many datasets, 16 GB for larger ones) and warns that on CPU only moderate datasets, up to about 5,000 samples, are feasible. The package ships the base and Fast checkpoints; Plus and Thinking are served through Prior Labs' API, and the company lists availability on AWS SageMaker and SAP AI Core. A cloud client and a set of community extensions sit alongside the core package.

Limits and Licence

Because the training set is the model's context, memory and prediction latency scale with the number of rows, and the Nature paper acknowledges that inference can be slower than a highly optimised tree library such as CatBoost. TabPFN-2.5 added distillation into a compact neural network or tree ensemble for low-latency serving.

The licence is the constraint most often overlooked. The TabPFN-3.5 weights are published under a licence that permits research, testing and internal evaluation but states that the model, its derivatives and its outputs cannot be used for any commercial or production purpose without a separate enterprise licence from Prior Labs. Only TabPFN v2 remains usable under Apache-style terms. Teams that need a permissively licensed in-context model should look at the alternatives listed under tabular foundation models; the head-to-head trade-offs with trees are in TabPFN vs XGBoost.

Further Reading