Method · Step 3

Why we validate out-of-sample

A revenue forecast is a claim about money that does not exist yet. On its own it is unfalsifiable, which is another way of saying worthless. It becomes trustworthy only when it is measured against stores whose real takings you already know. That measurement is the whole of this step, and it is the step most tools quietly skip.

Calibrate to your own estate

We start from your own stores, not a generic benchmark. Your estate already encodes what your brand is worth in the market: your average spend, your draw, the way your format performs against a given demand mix. IRIS calibrates to that reality, so the forecast speaks in your numbers rather than an industry average that may have nothing to do with how your stores actually trade.

Calibrating a model to data and then judging it on that same data is easy and dishonest. A model can fit the stores it has seen almost perfectly and still be useless on the next site, because it has memorised rather than learned. That is where out-of-sample testing comes in.

Hold stores out, then predict them

The test is simple to state. We hold a set of your stores out of the calibration entirely. The model never sees their real revenue. Then we ask it to forecast them, and we compare its predictions against what those stores actually earned. Because the model was blind to the answers, the comparison is honest rather than flattering.

Plot predicted against actual for every held-out store and the picture is immediate. Points sitting tight to the diagonal mean the forecast tracks reality; points scattered wide mean it does not. From that scatter we read the average deviation, the typical gap between predicted and actual, which becomes the evidence behind the interval you will see on the forecast itself. This is the same predicted-versus-actual check shown in how IRIS forecasts a site.

Why most tools skip this

Out-of-sample validation is uncomfortable to publish, because it puts a number on how wrong the tool can be. It is far easier to show a confident point estimate and say nothing about how it was tested. We take the opposite view: the deviation band is not an embarrassment to hide, it is the credential. A tool that will not tell you its error is asking you to trust it blindly.

Some error is irreducible

Honesty runs both ways. Even a well-validated forecast will not be exactly right, and not only because data is imperfect. Between the day you decide and the day you open, the world moves: a competitor arrives, interest rates shift, a street reroutes, a downturn arrives. No model can see the macro conditions of a year that has not happened yet. Some error is structural, not a flaw to be engineered away. We report it rather than pretend it is not there.

Validation is where a forecast stops being a claim and starts being evidence. It is also what lets us attach an honest range to the next site, which is what the 80% interval is for.