Skip to main content

Command Palette

Search for a command to run...

A Forecasting Model That Correctly Refuses to Forecast

Updated
•5 min read•View as Markdown
M
Full stack engineer. 25+ years of production software engineering across e-commerce, workflow automation and agentic ai

It beat the baseline by 98.7% on one series and lost to repeating the last value on another. That contrast is the most useful single fact about it.


Part one covered what TimesFM 3.0 is: a stateless numeric forecaster, not anything you can instruct. Before pointing it at a market, I wanted a calibration baseline. Same model, same settings, held-out backtest, two synthetic series with known properties.

Series Versus naive baseline Direction accuracy
Seasonal with drift 98.7% better 100%
Random walk 9.8% worse 47%

The naive baseline is the crudest forecast available: assume the next value equals the last one.

Read those two rows together, because separately each one misleads.

The first row is real capability

A 98.7% improvement over naive with perfect direction accuracy is not a marginal win. The model captured the seasonality and the trend and extrapolated both cleanly, zero-shot, with no fitting to that series at all. It had never seen it.

That is the capability being sold, and it is genuinely there.

The second row is the honest part

On a random walk it performed worse than doing nothing, and its direction calls landed at 47%, which is a coin flip with a rounding error.

The instinct is to read that as failure. It is the opposite. A random walk is unpredictable by construction. There is no structure to find. The next step is independent of every step before it.

A model that appeared to forecast a random walk would be broken, not impressive. It would be reporting structure that does not exist, which in a backtest looks like skill and in production looks like losing money. The correct behaviour on an unpredictable series is to fail, and ideally to fail about as badly as the naive baseline.

So the second row is a model declining to hallucinate. That deserves credit, and it is the most flattering framing available for everything that follows.

Why this is the useful fact

Those two rows are a map. They tell you precisely which side of a line any given problem sits on.

Where it works: series with stable repeating structure. Demand forecasting. Energy load. Web traffic. Call volume. Inventory draw. Capacity planning. Sensor telemetry. Domains where next week broadly resembles last week plus a trend.

The commercial argument there is strong and it is not really about accuracy. It is that you do not have to fit a model per series. Point it at ten thousand products and get ten thousand forecasts, zero-shot, in the time it takes to read this paragraph. Anyone who has maintained a per-SKU forecasting pipeline understands why that matters.

Where it does not work: anything closer to a random walk than to a seasonal series.

Markets are the adversarial case, and structurally so

This is the part worth sitting with before reading the results in the next post.

A liquid market is not merely a hard forecasting problem. It is a problem that actively destroys the kind of structure this model is built to find.

Any stable, simply extrapolable pattern in a liquid instrument gets arbitraged away. That is not a market being difficult, it is the definition of a market working. Participants are paid to find exactly the patterns a seasonal extrapolator would find, and their finding them is what removes them.

What survives is conditional and regime-dependent. Structure that holds under one set of circumstances and inverts under another. Relationships that only appear after a particular kind of move. These are precisely the patterns that require the conditional logic and state that part one established the architecture cannot represent.

So the calibration table predicts the outcome. Market tape is going to sit on the random-walk side of that line at short horizons, and the model is going to behave the way it behaved on row two.

That is exactly what happened, and the next post has the numbers.

One caution against over-reading this

It would be convenient to conclude “markets are random walks, so nothing could have worked.” That conclusion is wrong and I want to block it now, because the rest of this series depends on not making it.

Low predictability by a linear measure does not mean unpredictable. Conditional structure can be strong while every simple statistic says noise. This project found some, by a method that had nothing to do with the model, and that result is part five.

The claim here is narrower: this data sits on the wrong side of this model’s line. Not that the data is empty.


Next: It Lost to Repeat-the-Last-Value. The actual financial results on 118,857 bars of futures tape, including a reframing designed specifically to play to the model’s strengths, which also failed.