TabPFN vs XGBoost

Comparison

TabPFN and XGBoost represent two generations of prediction on tabular data. XGBoost is a gradient-boosted decision tree library that fits a new model to each dataset. TabPFN is one of the tabular foundation models: a transformer pretrained on synthetic datasets that takes a labelled training table as context and predicts the test rows directly, without dataset-specific training.

On small and medium tables, independent benchmarking now favours TabPFN-class models on accuracy, often without tuning. XGBoost keeps clear advantages in scale, licensing, deployment simplicity and inference cost. The largest recent claims for TabPFN come from its developer, Prior Labs, and should be read as vendor-reported.

Feature Comparison

DimensionTabPFNXGBoost
Model typePretrained transformer performing in-context learning over the training tableGradient-boosted decision tree ensemble trained per dataset
OriginHollmann, Müller, Eggensperger and Hutter, arXiv July 2022; v2 in Nature, January 2025; developed by Prior LabsChen and Guestrin, arXiv March 2016
Current version (October 2026)TabPFN-3.5, technical report dated 15 September 20263.4.1, released 14 August 2026
Fitting a new datasetNo gradient training; predictions come from a forward pass over the training dataIterative tree building
Hyperparameter tuningThe 2022 paper describes the model as needing noneUsually tuned; TabArena scores it at default, tuned, and tuned-plus-ensembled settings
Stated data limitsTabPFN-3.5: up to 1,000,000 rows and 20,000 features (repository README); v2: 10,000 samples and 500 features"Beyond billions of examples" with distributed and external-memory training
HardwareGPU recommended (about 8 GB VRAM; 16 GB for some large datasets); CPU limited to about 5,000 samplesRuns on CPU; GPU and distributed training (Dask, Spark, Ray, Kubernetes) supported
LicenceCode Apache 2.0; TabPFN-2.5 through 3.5 weights under non-commercial licences; v2 weights Apache 2.0 with attribution; commercial enterprise licence offeredApache 2.0
Independent benchmark (TabArena, NeurIPS 2025)TabPFNv2 outperformed other approaches "by a large margin" on the 33 of 51 datasets within its limitsGradient-boosted trees "still strong contenders"; deep learning methods caught up under larger time budgets with ensembling
Vendor-reported resultsTabPFN-2.5: 100% win rate against default XGBoost on classification datasets up to 10,000 rows and 500 features, 87% up to 100K rows; TabPFN-3.5: first place on TabArenaNot applicable
Tuned XGBoost head-to-head for TabPFN-3.5Not publicly documented in the 3.5 technical reportNot publicly documented
Inference costNot optimized for real-time inference; may be slower than tuned tree libraries (Nature paper); distillation to an MLP or tree ensemble introduced with 2.5Lightweight tree evaluation; standard latency comparisons with TabPFN-3.5 are not publicly documented

Detailed Analysis

Two different kinds of model

XGBoost, described by Chen and Guestrin in 2016, builds trees sequentially, each correcting the errors of the ensemble so far, with engineering for sparse data, approximate split finding and out-of-core computation. The fitted model is specific to one dataset. Its documentation states that missing values are handled by learning a default branch direction, and that the library scales through distributed frameworks when data exceeds memory.

TabPFN inverts the workflow. The expensive learning happens once, at pretraining, on a large collection of synthetic tabular tasks. At use time the labelled rows are the prompt. The 2022 paper claimed classification on small datasets "in less than a second" with no hyperparameter tuning. The 2025 Nature paper reported that, on datasets up to 10,000 samples and 500 features, TabPFN in 2.8 seconds outperformed an ensemble of the strongest baselines tuned for four hours.

What independent benchmarks say

For years the consensus ran the other way. Grinsztajn, Oyallon and Varoquaux (2022) found tree-based models remained state of the art on medium-sized tabular data of about 10,000 samples, and attributed it to inductive biases neural networks lacked. TabArena, a continuously maintained benchmark of 51 curated datasets (Erickson et al., NeurIPS 2025), updated the picture: gradient-boosted trees "are still strong contenders", deep learning methods caught up under larger time budgets with ensembling, and "foundation models excel on smaller datasets". At launch TabPFNv2 could be run on only 33 of the 51 datasets because of its size limits, and within that subset it led by a wide margin.

Two notes on independence. TabArena's authors include Frank Hutter, a TabPFN co-author, although the benchmark is a multi-institution effort with public code. And the current live leaderboard could not be read for this page, so standings after the launch paper are taken from Prior Labs' own reports.

Vendor-reported progress since 2025

Prior Labs' TabPFN-2.5 report (November 2025) claims the top TabArena position, a 100% win rate against default XGBoost on small-to-medium classification datasets and 87% on datasets up to 100,000 rows. The comparison is with default, not tuned, XGBoost. The TabPFN-3.5 technical report (September 2026) claims first place on TabArena and gains of about 150 Elo on a broader vendor benchmark, with variants for text-rich columns, maximum accuracy and speed. That report gives no direct numerical comparison with tuned XGBoost, LightGBM or CatBoost, and its claims have not, as far as could be found, been independently reproduced.

Scale, cost and licensing

TabPFN's stated envelope has grown from 1,000 training points in 2022 to 1,000,000 rows in the 3.5 README. The published evidence at the top of that range is thinner than at the bottom. Because the training table is processed at prediction time, the Nature paper notes that memory scales with dataset size and that the model is not optimized for real-time inference. Prior Labs addresses this with a distillation engine that converts a fitted TabPFN into a compact MLP or tree ensemble, and with a faster 3.5 variant reported as up to six times quicker than the base model.

Licensing is a practical difference. XGBoost is Apache 2.0 throughout. The TabPFN code is Apache 2.0, but the repository states that weights for versions 2.5 through 3.5 are released under non-commercial licences, with commercial use available through an enterprise licence or the hosted API. Only the older v2 weights carry a permissive licence with an attribution requirement.

Best For

Small or medium table, accuracy matters, little time to tune

TabPFN

Foundation models lead TabArena on smaller datasets, and TabPFN needs no hyperparameter search.

Hundreds of millions of rows or distributed training

XGBoost

XGBoost is built for out-of-core and cluster training; TabPFN's stated ceiling is 1,000,000 rows.

Commercial production use with open-source licensing only

XGBoost

XGBoost is Apache 2.0; current TabPFN weights are non-commercial unless licensed.

Rapid baseline during exploration

TabPFN

A single fit-and-predict call gives a strong reference score without tuning.

Low-latency, high-volume online scoring on CPU

XGBoost

Tree inference is lightweight; TabPFN was described by its authors as not optimized for real-time use, though distillation narrows the gap.

No GPU available

XGBoost

TabPFN on CPU is limited to about 5,000 samples; XGBoost trains on CPU at scale.

Highest possible accuracy on a mid-sized dataset

Both / depends

TabArena found ensembles across model families advance the state of the art; combining both is the stronger approach.

Text-rich or high-cardinality columns

Both / depends

Prior Labs reports large gains for TabPFN-3.5 here, but the evidence is vendor-reported and has no tuned-tree head-to-head.

The Bottom Line

For tables of up to tens of thousands of rows, TabPFN is now a serious first choice on accuracy and effort: the independent TabArena study and the peer-reviewed Nature paper both support it within those limits. For very large data, strict latency, CPU-only environments or permissive licensing, XGBoost remains the safer default, and it is still competitive on accuracy when tuned.

As of October 2026 the open question is the middle: datasets of hundreds of thousands to a million rows, where TabPFN's reach is recent and the supporting results are mostly the vendor's. The practical course is to run both on the actual data with a proper validation split, compare against a tuned tree model and not a default one, and consider ensembling them. That is the configuration TabArena found strongest.