NVIDIA Kumo Tabular is an open foundation model for tabular data prediction, available on Hugging Face under the OpenMDW-1.1 license. It predicts labels for new rows in a single forward pass without training, tuning, or feature engineering, using a Transformer architecture with column, row, and in-context attention. The model, pretrained on artificial data in three sizes (28M–215M parameters), ranks first on four benchmarks and addresses enterprise machine learning tasks historically dominated by gradient-boosted trees.
From 2018 to 2022, language models exhibited emergent capabilities including natural language understanding, translation, arithmetic, code generation, and instruction-following. These abilities arose unexpectedly without task-specific training, with performance improving dramatically as model scale increased.
TabPFN and TabICL, tabular foundation models pretrained on synthetic data, outperformed tuned XGBoost on all 14 datasets from the Grinsztajn benchmark without requiring training on new tables. These models use in-context learning to make predictions in a single forward pass, potentially eliminating the need for hyperparameter tuning that traditionally consumes significant computational resources.