Back to News & insightsGuides

Tabular AI: choose a model for the rows you actually have

Compare trees and neural networks on honest data splits.

Editorial guide · Updated September 28, 2026 · 8 min read
A branching silver tree and woven metal mesh rise from an ordered grid of graphite blocks.

Many useful AI problems begin in an ordinary table. Each row records an event, an item, or a person, while columns describe quantities and categories. There may be no paragraph to generate and no photograph to interpret. The task could be estimating how long a repair will take or identifying which inventory items need closer inspection. These problems deserve their own modeling choices rather than inheriting whatever architecture dominates AI headlines.

Tabular modeling is challenging because the table's convenient appearance hides its history. Columns may be recorded at different times, missing values may have several meanings, and rows may repeatedly describe the same underlying entity. Before comparing neural networks with decision trees, establish what a prediction is allowed to know. An honest data boundary usually matters more than an impressive model name.

Define the moment of prediction

Consider a hypothetical bicycle-repair shop that wants to estimate turnaround time when a customer checks in a bicycle. The available columns include the reported problem, bicycle type, current workshop queue, and parts already in stock. The final invoice and technician's completion notes appear later.

If those later columns enter training as features, the model receives clues that will not exist when the estimate is needed. A random train-test split does not fix that problem. The information boundary must be enforced when the training examples are constructed, using what was actually available at check-in.

Write a small example row by hand and label each field with its availability time. This exercise often reveals ambiguous columns. A status called parts available might mean available at arrival or available after an order was delivered. The name alone is not a reliable description of the evidence.

Establish simple baselines

Start with a prediction that can be explained without a complex model. The shop might use the typical turnaround for a repair category or a simple relationship between queue length and completion time. These baselines expose how much improvement a more complicated system must deliver to justify its maintenance.

A baseline also helps diagnose the target. If all methods struggle during unusual supplier delays, the missing information may matter more than model capacity. If even a crude rule performs suspiciously well, investigate leakage before celebrating. Simple models can function as probes into the dataset's construction.

Keep the evaluation metric connected to the decision. An average absolute error describes one aspect of an estimate, but consistently underestimating long repairs can create a different problem from small symmetric mistakes. Inspect the distribution of errors and the cases that affect customer communication most.

Why tree ensembles deserve a fair comparison

Decision trees divide the input space through a sequence of conditions. Ensembles combine many such components, offering flexible prediction without requiring every relationship to be smooth. For tables with mixed feature types and irregular interactions, this can be a useful starting point.

Léo Grinsztajn, Edouard Oyallon, and Gaël Varoquaux studied why tree-based models remained strong on a benchmark collection of tabular tasks. Their results support taking tree baselines seriously within that study's setting. They do not establish a permanent winner for every table, dataset size, or future architecture.

Read the original research paper on arXiv

For the repair shop, compare a tuned tree ensemble with the simple baseline under the same split and target definition. Do not assume that familiar business columns make the problem trivial. Queue pressure, parts availability, and repair category can interact in ways that still require careful validation.

Neural networks need an appropriate setup

Neural models can learn flexible representations of numerical and categorical inputs. Their performance depends on preprocessing, architecture, regularization, optimization, and the amount and nature of the data. A weak configuration is not a meaningful representative of everything neural modeling can do.

Yury Gorishniy and colleagues' study of deep learning for tabular data examines strong neural baselines, including a transformer-based approach. It is useful evidence that comparisons should specify implementations and tuning rather than treating either neural networks or trees as one fixed method.

Read the original research paper on arXiv

Give candidate methods reasonable, documented tuning budgets. An extensively searched tree model and a default neural model do not receive equivalent preparation. The reverse comparison is equally misleading. Track total development effort as well as final accuracy so the team can choose something it can realistically reproduce.

Missing values are part of the story

A missing measurement can mean not collected, not applicable, unavailable at prediction time, or lost during processing. Replacing all missing entries with one convenient value erases these distinctions. Sometimes the pattern of missingness is itself related to the workflow that produced the outcome.

In the shop example, an empty parts code could mean no replacement is needed or that diagnosis has not happened yet. Those meanings should not be silently merged if the source system can distinguish them. Improve the schema where possible and document unavoidable ambiguity.

Fit any imputation or scaling procedure using the training partition only. Otherwise information about the held-out data influences the pipeline before evaluation. Save preprocessing with the model so that a production request follows the same transformations as the examples used to establish quality.

Split rows according to the deployment question

If the shop wants to predict future repairs, a time-based evaluation is usually more informative than shuffling all years together. It exposes changes in staffing, suppliers, and customer demand. However, the exact split should follow the deployment question rather than a universal preference for one technique.

Repeated entities require additional thought. The same bicycle or customer may appear many times. If a product must work for entirely new customers, evaluate that separately from returning customers. Otherwise memorized patterns can make generalization appear stronger than it will be for the intended population.

Keep the final test set independent of repeated model selection. Use development partitions to choose settings, then evaluate the chosen pipeline on examples that did not guide those choices. If the team learns from the final failures and changes the model, it needs fresh evidence for the revised version.

Watch categorical shortcuts

An identifier can be useful when it represents a stable entity with meaningful history, but it can also become an accidental shortcut. A job number may encode time or location in a way that disappears when the numbering system changes. A staff identifier may absorb patterns that actually reflect different assignments.

Investigate unusually influential identifiers with targeted comparisons. Remove or alter them in a controlled experiment and inspect which predictions change. This does not mean every identifier must be excluded. It means the team should understand why the signal is expected to survive deployment.

Plan for categories that did not occur during training. A new bicycle brand or repair code should produce defined behavior instead of a preprocessing exception. Evaluation examples should include these cases, along with malformed or unexpectedly encoded inputs that real integrations can send.

Evaluate slices that affect the decision

A single average can conceal poor estimates for uncommon repairs. Break down errors by repair type, queue conditions, parts availability, and other operationally relevant factors. Small groups may have noisy estimates, so report their sample sizes and avoid pretending every observed difference is stable.

For turnaround estimates, inspect both ordinary cases and the long tail. A model can perform well on frequent quick repairs while missing the delays that most frustrate customers. Choose thresholds and communication rules with those consequences in mind, rather than optimizing a metric that ignores them.

Review examples with workshop staff. They may identify mislabeled outcomes or explain that a long completion time included a customer delaying collection. That discovery changes the meaning of the target. More elaborate modeling cannot repair a label that measures something different from the promise the product intends to make.

Treat uncertainty as a separate output requirement

A point estimate can be useful, but a single number may imply more precision than the evidence supports. If the shop needs a range, define what the range is intended to mean and evaluate its coverage on relevant held-out data. Do not manufacture a range by adding an arbitrary constant around every prediction.

Some conditions may justify requesting more information or using a conservative manual estimate. Record these decisions and evaluate how often they occur. A system that recognizes an unfamiliar repair can be more useful than one that always produces a precise-looking answer.

The interface should describe the estimate in terms the staff and customer can use. It need not expose the model family. What matters is whether the output supports a clear next step and whether people can understand when the estimate is uncertain or no longer applicable.

Choose the pipeline the team can sustain

The final comparison should include prediction quality, startup and serving cost, retraining effort, schema stability, and recovery from broken inputs. A small improvement may be worthwhile, but it should be weighed against the operational work required to preserve it as the table changes.

Save the feature definitions, preprocessing, split logic, and evaluation results with the chosen artifact. Monitor whether new data still matches the assumptions used to approve it. Tabular AI succeeds when a disciplined representation of the problem meets a suitable model, and the resulting prediction remains honest about what was known at the moment it was made.

Sources and rights

The two linked research papers are available under arXiv's non-exclusive distribution licence, with copyright retained by their authors. This original explainer does not reproduce their tables, figures, code, or prose. The repair-shop scenario and evaluation discussion are illustrative.

An original editorial guide. Provider capabilities and documentation can change. Follow the linked sources and test the exact model or service before relying on it.