Time-series foundation models: forecast the future without leaking it
Design realistic forecasting evaluations with historical cutoffs, operational baselines, uncertainty, missing data, and a clear connection to the decision being made.

Forecasting looks straightforward when the future is already in the spreadsheet. A model receives a historical sequence, predicts what comes next, and earns a score against the values that eventually occurred. The difficult part is ensuring that the experiment gives the model only information that would actually have existed at the time of the decision.
Time-series foundation models make forecasting more accessible by offering pretrained systems that can be applied to new sequences. That convenience is useful, but it does not remove the need to define the forecast, construct honest historical tests, and connect errors to operational consequences. A promising prediction curve is the beginning of an investigation, not the conclusion.
What is being transferred from pretraining
The original TimesFM research explores a pretrained decoder-based model for time-series forecasting. Chronos takes another approach, representing scaled numerical values through tokenization and training probabilistic forecasting models. These papers establish examples of learning across many series; they do not establish a universal winner for every future dataset.
Read the original research paper on arXiv
Read the original research paper on arXiv
The practical attraction is a starting point that may need less task-specific training. However, a pretrained system still receives a particular representation of your series and operates under particular assumptions. Check the exact implementation's treatment of frequency, missing values, context, covariates, and forecast horizon before treating it as a drop-in component.
Define the decision before defining the metric
Consider a hypothetical community library that plans staffing for the next seven days. Its target is daily visitor count. The operational decision is how many staff hours to schedule, and the cost of understaffing differs from the cost of a quiet shift with extra coverage. This is not a financial prediction or a reported deployment.
Specify when the schedule must be finalized. If the decision is made each Thursday morning, the forecast cannot use visitor counts from later that Thursday. Decide whether the target is the full building, individual branches, or service desks. Aggregating everything may hide precisely the variation that determines staffing needs.
Also define what the forecast does not decide. A predicted busy day might prompt a manager to review the schedule, rather than automatically changing workers' hours. A clear decision boundary makes it easier to evaluate whether a more accurate prediction is valuable enough to justify additional complexity.
Reconstruct what was known at each cutoff
A historical test should behave like a sequence of past decisions. At each cutoff, assemble the data that would have been available then, produce the forecast, and compare it with the later outcome. Move the cutoff forward and repeat. This is more informative than one convenient split selected after inspecting the results.
Availability time matters as much as event time. A visitor-count correction entered two weeks later should not appear in the earlier input unless the real system would have had access to it. Keep the distinction explicit in data preparation. Otherwise a careful-looking backtest can quietly benefit from hindsight.
The same rule applies to external variables. A planned event known in advance may be legitimate input. An event attendance total measured afterward is not. For weather, a historical observed temperature and a weather forecast issued before the staffing decision are different inputs with different information content.
Give simple baselines a fair chance
Begin with a seasonal baseline, such as using the corresponding day from a recent week, and a simple historical average appropriate to the calendar. These baselines are not embarrassing competitors. They help show whether the problem contains predictable structure and whether the more complex system captures anything operationally useful beyond it.
Evaluate all candidates on the same cutoffs and targets. Do not give the foundation model cleaned data while leaving the baseline with untreated gaps. Likewise, count any manual tuning or exception rules added after viewing results. A comparison is only meaningful if the development effort and available information are visible.
If a baseline performs similarly, inspect where the models differ. The more complex system might help on unusual event days while adding little on routine days. That finding could support a targeted workflow rather than replacing every forecast. It could also reveal that the apparent improvement comes from information unavailable at decision time.
Missingness is a feature of the data-generating process
A zero can mean nobody visited, the building was closed, or the counting device failed. Those meanings should not be merged automatically. A model trained on an ambiguous sequence may learn patterns that reflect instrumentation problems rather than real demand. Maintain a separate explanation for gaps and exceptional observations.
Create a documented preparation policy. Identify which gaps can be filled, which periods should be excluded, and which require an explicit indicator if the implementation supports one. Test the policy using only past information. Interpolating a missing point from later observations can create leakage in a historical forecast experiment.
Keep raw and prepared data separately. When a forecast behaves strangely, reviewers should be able to determine whether the cause was the original record, the preparation policy, or the model. Reproducible preprocessing is part of the forecasting system, even when the prediction itself comes from an external service.
Measure the mistakes people actually care about
Report an error measure in the target's natural units where possible, alongside any normalized comparison needed across branches. A single percentage can become unstable or confusing when actual visitor counts are near zero. Explain the metric so a manager can connect it to a staffing decision.
Separate systematic overprediction from underprediction. Two models with similar average absolute error can behave differently if one persistently underestimates busy periods. Inspect results by day of week, branch, horizon, and unusual-event status. These slices should be planned around the operating problem rather than selected only because they flatter one model.
Evaluate the full horizon. A system that performs well tomorrow may be much less useful six days ahead, when staffing is harder to change. Present the horizon-specific evidence clearly. Averaging across all days can hide the weakness at exactly the lead time the organization needs.
Uncertainty must survive contact with new conditions
If the model produces predictive ranges, inspect both their coverage and their width on held-out historical periods. A range that contains nearly everything but is too wide to guide a decision may have limited value. A narrow range that misses busy days can create unjustified confidence.
Do not treat an interval as protection against every possible change. A branch renovation, a new counting device, or a changed opening schedule may make historical patterns less relevant. Record these changes and evaluate whether the forecast process should pause, adapt, or receive additional review.
The interface should explain the meaning of the displayed uncertainty. A manager needs to know whether a range describes model variability, an empirically checked predictive interval, or something else. Avoid presenting a shaded band simply because it looks scientific. Its interpretation should be tied to the method actually used.
Work through a disagreement between forecasts
Suppose the seasonal baseline predicts a quiet Saturday while the pretrained model predicts a busy one. Before choosing the more elaborate model, inspect what information each received. Was a public event included? Did a closure in the previous week distort the baseline? Was the model given a corrected count that would not have been available when the schedule was set?
If both forecasts used legitimate information, record which operational decision each would support and compare the later outcome. One weekend cannot establish superiority, but it can reveal a useful test category. Add comparable event weekends to the historical evaluation using the same cutoff rules, rather than changing the test only for the model that produced the preferred answer.
Also examine the cost of the disagreement. If both forecasts lead to the same staffing plan, their numerical difference may have little immediate value. If they lead to different plans, the error tradeoff deserves closer review. This connects forecasting improvements to decisions without pretending that every smaller numerical error creates an equally valuable operational gain.
Run a shadow period before changing operations
For several decision cycles, generate forecasts without letting them control staffing. Compare the recommendation with the existing process and record whether it would have changed a decision. Ask managers to identify missing contextual information, such as an event that never entered the data pipeline.
Use the shadow period to test practical failures: delayed data, unavailable model endpoints, malformed outputs, and unexpected calendar entries. Define a fallback forecast that can be produced reliably. A slightly less accurate forecast delivered on time may be more useful than a sophisticated result that arrives after the schedule is finalized.
Adopt the model only with a clear account of the improvement, its limits, and the conditions that trigger reevaluation. Time-series foundation models can reduce the effort required to try forecasting ideas. The durable advantage comes from building a trustworthy decision process around them, with historical honesty, understandable errors, and operational ownership.