Build1 publisher3 min readPublished
NVIDIA's Kumo Tabular predicts labels for new table rows in a single forward pass
NVIDIA released Kumo Tabular, a commercially licensed tabular model it says ranks first on four public benchmarks with no per-task training. At inference the labeled table replaces the fitted model, so the rankings carry over only where a team's tables resemble the benchmark sets.
The Engineer · Build desk

What happened
- The model was pretrained only on artificial data and ships in three sizes, from 28 million to 215 million parameters.
- It is a Transformer that alternates column and row attention before a final in-context stage, following designs introduced in TabICL and TabPFN.
- Classification returns class probabilities, while regression returns 999 quantiles from which a point prediction and an uncertainty estimate follow.
- Weights are on Hugging Face and code is in NVIDIA's structured-data-models repository on GitHub, under the OpenMDW-1.1 license.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision A team with a working tree pipeline can get a head-to-head on its own held-out split for the cost of inference, because there is no fit or hyperparameter search to configure first.
- cost Serving cost follows the labeled table: a bigger context means a larger cache to hold and read on every prediction, so label volume sets the memory budget.
- constraint The benchmark rankings say least about very large labeled tables, the case where synthetic pretraining left a gap NVIDIA had to close with a length-dependent attention temperature.
A call to Kumo Tabular takes two inputs: a table of labeled rows and the rows to be scored. Nothing is fitted in between. It returns class probabilities or numeric predictions from one forward pass [2]. NVIDIA calls this in-context learning for tables, with a model pretrained on millions of tables reading the labeled table as its context [15].
Each attention stage has a specific job [6]. Cells become tokens through Fourier features, with separate weights for numerical and categorical values, and missing values go in without imputation [7]. Column attention learns whether a value such as 42 is typical or extreme for its column. It uses induced self-attention, so its cost grows linearly with the number of rows [8]. Row attention learns how features interact, and four learnable CLS tokens then compress each row. From that point on, cost no longer depends on the number of columns [9].
The final stage is the part I would copy into my own designs. Context rows attend to each other and query rows attend only to context rows, so a prediction depends on the context and the row itself, not on which other rows share its batch [10]. Because the context never sees the queries, its keys and values are computed once and reused, and Test-GQA shrinks the cache each prediction reads [11]. A service can hold one labeled table's cache and stream new rows against it [11].
One design choice points at the constraint behind the whole model. Pretraining used only artificial data [3]. NVIDIA writes that softmax attention spreads out as the number of keys grows. Attention that is sharp over a few hundred rows can dissolve over tens of thousands, the case when an inference table is much larger than a typical training table [13]. The fix is a query temperature that grows with the logarithm of length [13]. A production table with tens of thousands of labeled rows sits in the regime that needed it [13].
NVIDIA says the model ranks first on TabArena, BeyondArena, TALENT and ScoringBench [5]. The announcement states the ranking without margins over tree baselines. For the result to carry over, a team's tables would need row and column counts inside the range those suites cover. Their feature distributions would need to resemble what the synthetic generator produced [3]. And the serving budget would have to allow every query to attend over the full context [10].
NVIDIA sets the model against two decades of gradient-boosted trees, where every new question means collecting labels, engineering features, searching hyperparameters, validating and deploying [14]. Kumo Tabular removes the feature work, the search and the training. Labels are still needed as context [2]. The largest checkpoint has about 7.7 times the parameters of the smallest [1], so a first comparison on the small one is cheap. My context is a team with a tuned tree model and a fixed held-out split. There, the test is one inference run per checkpoint [2]. I think that is cheap enough to run before building the next tree pipeline, and the replacement decision belongs to that split's numbers.
What to watch
- Independent reruns of the TabArena, TALENT or ScoringBench results that publish Kumo Tabular's margin over tuned gradient-boosted trees.
- Latency and memory figures for the 215M checkpoint with a context of tens of thousands of labeled rows.
- Accuracy reports on production tables far larger than the synthetic pretraining tables, where the length-aware temperature applies.