TabPFN vs XGBoost
ComparisonTabPFN and XGBoost represent two generations of prediction on tabular data. XGBoost is a gradient-boosted decision tree library that fits a new model to each dataset. TabPFN is one of the tabular foundation models: a transformer pretrained on synthetic datasets that takes a labelled training table as context and predicts the test rows directly, without dataset-specific training.
On small and medium tables, independent benchmarking now favours TabPFN-class models on accuracy, often without tuning. XGBoost keeps clear advantages in scale, licensing, deployment simplicity and inference cost. The largest recent claims for TabPFN come from its developer, Prior Labs, and should be read as vendor-reported.
Feature Comparison
| Dimension | TabPFN | XGBoost |
|---|---|---|
| Model type | Pretrained transformer performing in-context learning over the training table | Gradient-boosted decision tree ensemble trained per dataset |
| Origin | Hollmann, Müller, Eggensperger and Hutter, arXiv July 2022; v2 in Nature, January 2025; developed by Prior Labs | Chen and Guestrin, arXiv March 2016 |
| Current version (October 2026) | TabPFN-3.5, technical report dated 15 September 2026 | 3.4.1, released 14 August 2026 |
| Fitting a new dataset | No gradient training; predictions come from a forward pass over the training data | Iterative tree building |
| Hyperparameter tuning | The 2022 paper describes the model as needing none | Usually tuned; TabArena scores it at default, tuned, and tuned-plus-ensembled settings |
| Stated data limits | TabPFN-3.5: up to 1,000,000 rows and 20,000 features (repository README); v2: 10,000 samples and 500 features | "Beyond billions of examples" with distributed and external-memory training |
| Hardware | GPU recommended (about 8 GB VRAM; 16 GB for some large datasets); CPU limited to about 5,000 samples | Runs on CPU; GPU and distributed training (Dask, Spark, Ray, Kubernetes) supported |
| Licence | Code Apache 2.0; TabPFN-2.5 through 3.5 weights under non-commercial licences; v2 weights Apache 2.0 with attribution; commercial enterprise licence offered | Apache 2.0 |
| Independent benchmark (TabArena, NeurIPS 2025) | TabPFNv2 outperformed other approaches "by a large margin" on the 33 of 51 datasets within its limits | Gradient-boosted trees "still strong contenders"; deep learning methods caught up under larger time budgets with ensembling |
| Vendor-reported results | TabPFN-2.5: 100% win rate against default XGBoost on classification datasets up to 10,000 rows and 500 features, 87% up to 100K rows; TabPFN-3.5: first place on TabArena | Not applicable |
| Tuned XGBoost head-to-head for TabPFN-3.5 | Not publicly documented in the 3.5 technical report | Not publicly documented |
| Inference cost | Not optimized for real-time inference; may be slower than tuned tree libraries (Nature paper); distillation to an MLP or tree ensemble introduced with 2.5 | Lightweight tree evaluation; standard latency comparisons with TabPFN-3.5 are not publicly documented |
Detailed Analysis
Two different kinds of model
XGBoost, described by Chen and Guestrin in 2016, builds trees sequentially, each correcting the errors of the ensemble so far, with engineering for sparse data, approximate split finding and out-of-core computation. The fitted model is specific to one dataset. Its documentation states that missing values are handled by learning a default branch direction, and that the library scales through distributed frameworks when data exceeds memory.
TabPFN inverts the workflow. The expensive learning happens once, at pretraining, on a large collection of synthetic tabular tasks. At use time the labelled rows are the prompt. The 2022 paper claimed classification on small datasets "in less than a second" with no hyperparameter tuning. The 2025 Nature paper reported that, on datasets up to 10,000 samples and 500 features, TabPFN in 2.8 seconds outperformed an ensemble of the strongest baselines tuned for four hours.
What independent benchmarks say
For years the consensus ran the other way. Grinsztajn, Oyallon and Varoquaux (2022) found tree-based models remained state of the art on medium-sized tabular data of about 10,000 samples, and attributed it to inductive biases neural networks lacked. TabArena, a continuously maintained benchmark of 51 curated datasets (Erickson et al., NeurIPS 2025), updated the picture: gradient-boosted trees "are still strong contenders", deep learning methods caught up under larger time budgets with ensembling, and "foundation models excel on smaller datasets". At launch TabPFNv2 could be run on only 33 of the 51 datasets because of its size limits, and within that subset it led by a wide margin.
Two notes on independence. TabArena's authors include Frank Hutter, a TabPFN co-author, although the benchmark is a multi-institution effort with public code. And the current live leaderboard could not be read for this page, so standings after the launch paper are taken from Prior Labs' own reports.
Vendor-reported progress since 2025
Prior Labs' TabPFN-2.5 report (November 2025) claims the top TabArena position, a 100% win rate against default XGBoost on small-to-medium classification datasets and 87% on datasets up to 100,000 rows. The comparison is with default, not tuned, XGBoost. The TabPFN-3.5 technical report (September 2026) claims first place on TabArena and gains of about 150 Elo on a broader vendor benchmark, with variants for text-rich columns, maximum accuracy and speed. That report gives no direct numerical comparison with tuned XGBoost, LightGBM or CatBoost, and its claims have not, as far as could be found, been independently reproduced.
Scale, cost and licensing
TabPFN's stated envelope has grown from 1,000 training points in 2022 to 1,000,000 rows in the 3.5 README. The published evidence at the top of that range is thinner than at the bottom. Because the training table is processed at prediction time, the Nature paper notes that memory scales with dataset size and that the model is not optimized for real-time inference. Prior Labs addresses this with a distillation engine that converts a fitted TabPFN into a compact MLP or tree ensemble, and with a faster 3.5 variant reported as up to six times quicker than the base model.
Licensing is a practical difference. XGBoost is Apache 2.0 throughout. The TabPFN code is Apache 2.0, but the repository states that weights for versions 2.5 through 3.5 are released under non-commercial licences, with commercial use available through an enterprise licence or the hosted API. Only the older v2 weights carry a permissive licence with an attribution requirement.
Best For
Small or medium table, accuracy matters, little time to tune
TabPFNFoundation models lead TabArena on smaller datasets, and TabPFN needs no hyperparameter search.
Hundreds of millions of rows or distributed training
XGBoostXGBoost is built for out-of-core and cluster training; TabPFN's stated ceiling is 1,000,000 rows.
Commercial production use with open-source licensing only
XGBoostXGBoost is Apache 2.0; current TabPFN weights are non-commercial unless licensed.
Rapid baseline during exploration
TabPFNA single fit-and-predict call gives a strong reference score without tuning.
Low-latency, high-volume online scoring on CPU
XGBoostTree inference is lightweight; TabPFN was described by its authors as not optimized for real-time use, though distillation narrows the gap.
No GPU available
XGBoostTabPFN on CPU is limited to about 5,000 samples; XGBoost trains on CPU at scale.
Highest possible accuracy on a mid-sized dataset
Both / dependsTabArena found ensembles across model families advance the state of the art; combining both is the stronger approach.
Text-rich or high-cardinality columns
Both / dependsPrior Labs reports large gains for TabPFN-3.5 here, but the evidence is vendor-reported and has no tuned-tree head-to-head.
The Bottom Line
For tables of up to tens of thousands of rows, TabPFN is now a serious first choice on accuracy and effort: the independent TabArena study and the peer-reviewed Nature paper both support it within those limits. For very large data, strict latency, CPU-only environments or permissive licensing, XGBoost remains the safer default, and it is still competitive on accuracy when tuned.
As of October 2026 the open question is the middle: datasets of hundreds of thousands to a million rows, where TabPFN's reach is recent and the supporting results are mostly the vendor's. The practical course is to run both on the actual data with a proper validation split, compare against a tuned tree model and not a default one, and consider ensembling them. That is the configuration TabArena found strongest.
Further Reading
- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second (Hollmann et al., 2022) – arXiv
- Accurate predictions on small data with a tabular foundation model (Hollmann et al., January 2025) – Nature
- TabPFN-2.5: Advancing the State of the Art in Tabular Foundation Models (November 2025) – arXiv / Prior Labs
- TabPFN-3.5: Technical Report (September 2026) – Prior Labs
- TabPFN repository: limits, hardware and licensing – GitHub
- TabArena: A Living Benchmark for Machine Learning on Tabular Data (Erickson et al., NeurIPS 2025) – arXiv
- XGBoost: A Scalable Tree Boosting System (Chen and Guestrin, 2016) – arXiv
- XGBoost Documentation – xgboost.readthedocs.io
- Why do tree-based models still outperform deep learning on tabular data? (Grinsztajn et al., 2022) – arXiv