The Current Generation of Tabular Foundation Models: A Critical Review

Citation

Kurashkin, Sergei O. and Tynchenko, Vadim S. and Borodulin, Aleksei S. and Nelyub, Vladimir A. and Kalutsky, Nikolay O. and Connie, Tee (2026) The Current Generation of Tabular Foundation Models: A Critical Review. Machine Learning and Knowledge Extraction, 8 (8). p. 244. ISSN 2504-4990

[img] Text
make-08-00244-v2.pdf - Published Version
Restricted to Repository staff only

Download (823kB)

Abstract

Tabular foundation models (TFMs) have moved tabular machine learning from per-dataset training towards amortised in-context inference, fitting a small-to-medium table in a single forward pass without a training run. The 2024–2026 release train, the TabPFN and TabICL lines and challengers such as Mitra, LimiX and Orion, has produced a generation whose architectures, capabilities and limits are documented mainly in preprints, while existing surveys treat these models as a subsection of tabular deep learning or of language-model table understanding. This review is, to our knowledge, the first organised around the current generation. From a corpus of 961 screened records and 98 retained studies, it taxonomises the architectures by pretraining regime, maps the capability space across five axes, isolates the language-model-on-tabular strand for prediction, feature engineering and generation, and summarises openness and deployment. A dedicated critical synthesis then reads the reported capabilities against independent evidence: on the studies reviewed here, tree-based and deep models retain the lead across 142 curated datasets that go beyond the standard independent and identically distributed setting; on 112 datasets, the models attain the highest accuracy but weaker conditional coverage than gradient-boosted trees; and robustness under feature shift, fairness and generation quality remain open. Amortised in-context prediction is thus a working paradigm whose independent evidence has yet to match its benchmark claims.

Item Type: Article
Uncontrolled Keywords: tabular foundation models, in-context learning, prior-data fitted networks, TabPFN, tabular deep learning, large language models, synthetic data generation, benchmark evaluation, calibration, fairness
Subjects: Q Science > Q Science (General)
Divisions: Faculty of Information Science and Technology (FIST)
Depositing User: Ms Suzilawati Abu Samah
Date Deposited: 03 Sep 2026 02:08
Last Modified: 03 Sep 2026 02:08
URII: http://shdl.mmu.edu.my/id/eprint/16575

Downloads

Downloads per month over past year

View ItemEdit (login required)