Forecast Accuracy
Forecast accuracy is the measured gap between a frozen forecast and the demand that actually occurred, expressed through an error metric — MAPE, WAPE (weighted MAPE, or WMAPE) or bias — at a stated level of the product hierarchy and a stated horizon. A figure quoted without its metric, level and horizon cannot be compared with any other.
The metrics answer different questions. MAPE, mean absolute percentage error, averages |forecast − actual| ÷ actual period by period, so a one-unit week counts as much as a two-hundred-unit week, and it is undefined in any period that sold nothing. WAPE, weighted absolute percentage error, divides the sum of absolute errors by the sum of actuals: each miss is weighted by the volume behind it, and the figure stays defined at zero demand. WMAPE — MAPE weighted by actuals — reduces to the same calculation. Bias keeps the sign, the sum of forecast minus actual over the sum of actuals, and it is the only one of the three that says which way a forecast leans. An accuracy percentage quoted as 100% minus WAPE carries no direction either.
As an illustration, one style-color sells 40, 60, 100 and 200 units over four weeks against a forecast of 50, 50, 120 and 180. The absolute errors are 10, 10, 20 and 20, so WAPE is 60 ÷ 400 = 15.0%. MAPE is the mean of 25.0%, 16.7%, 20.0% and 10.0%, or 17.9% — higher, because the two smallest weeks carry the largest percentage misses. Bias is zero: the forecast totals exactly the 400 that sold. Now raise every week of the forecast by 10%. Bias becomes +10.0%, yet WAPE improves to 13.5%, because the lift closes the two under-forecast weeks by more than it widens the two over-forecast ones. The better WAPE belongs to a forecast that now calls 440 units against 400 sold — which is why magnitude and direction are reported as a pair, never one without the other.
Three choices decide whether the number means anything. The level: errors cancel as they aggregate, so the same forecast can only read the same or better at class by season than at style-color by week, and accuracy belongs at the grain the decision is taken — size by week for a replenishment trigger, style-color for a buy. The version: grade the forecast that was frozen when the decision was made, not the latest reforecast, which has already absorbed some of the actuals it is being graded against. And censored demand: in a week a size is out of stock, sales record what was on hand rather than what customers wanted, so a forecast can look most accurate exactly when service failed. Bias is the correctable part — adjust the model or the override that produces it. The error left once bias is removed is not forecast away; it is carried as safety stock or held back as open-to-buy reserve. The tracking signal that detects drift, and the reporting spec that makes a figure comparable to itself, are worked through in the forecast accuracy guide.
RetailNorthstar puts these metrics where planning decisions happen — connected to one plan, live against actuals.
Book a Demo →