Retail Planby RetailNorthstar

How to measure forecast accuracy

Forecast accuracy is only meaningful when the metric, the aggregation level and the horizon are all stated — a single accuracy number with none of them attached carries no information, because the same forecast against the same actuals produces wildly different figures depending on which of the three you choose. Measuring it well is not a modelling problem. It is a bookkeeping problem: freeze what you are grading, name the grain, name the horizon, and refuse to let the number improve for reasons that have nothing to do with the forecast.

This guide is about measurement, not method. It does not cover how to build a forecast, which models suit which demand pattern, or how often to rebuild one — those sit in demand forecasting for fashion and how often to reforecast a merchandise plan on RetailNorthstar. What follows is the arithmetic of grading one, and the specific ways an accuracy number gets better without the forecast getting better.

Covered below: MAPE, WAPE, bias and the tracking signal with their formulas and their failure modes; why bias and error are different failures needing different fixes; the three things that make an accuracy number meaningless; how stockouts censor demand and flatter the result; a reporting spec; seven verticals; and eight mistakes. Every figure in the worked example is illustrative.

The short version
An accuracy figure is not a measurement until it carries six things: the metric, the grain the error is summed at, the forecast horizon, the date the forecast was frozen, the availability rule, and the population it covers. Use WAPE for magnitude because it weights by the volume that costs money and survives zero demand; report bias beside it because direction determines the fix; run a tracking signal because drift is a different failure from size. Aggregating, reforecasting and stocking out all improve the number without improving the forecast.
Definition — Forecast accuracy
Forecast accuracy is the degree to which a frozen forecast matched what subsequently sold, measured at a stated grain and over a stated horizon. It is a family of measurements rather than one, because the choice of metric, level and horizon each change the answer on identical data. Expressed as one minus WAPE it is bounded above at 100% and unbounded below, which is why a negative accuracy figure is arithmetically ordinary rather than a data error.
Accuracy = 1 − WAPE, where WAPE = Σ|forecast − actual| ÷ Σactual
Used by: Demand planners, merchandise planners and buyers, pre-season and in-season
Related: MAPE, WAPE, forecast bias, tracking signal, mean absolute deviation, censored demand, aggregation level

Four numbers, one dataset, an eight-point spread

The four metrics below are computed on the same twelve weeks of the worked example. They are not variations on a theme. Each answers a different question, each fails in a different place, and the gap between the highest and lowest figure is larger than the difference most teams are trying to detect season over season.

MAPE, WAPE, bias and tracking signal compared: formula, value on the illustrative twelve-week series, and what each one does and does not tell you.
MetricFormulaOn this seriesWhat it does and does not say
MAPEmean of |F − A| ÷ A, taken period by period21.5% (with the zero-demand week deleted)Undefined the moment an actual is zero, and the standard fix is to delete that period — which here silently discards 35 of the 163 units of absolute error, a fifth of the total miss. Weights a one-unit week the same as a seventy-unit week, and punishes over-forecasting without limit while capping under-forecasting at 100%.
WAPEΣ|F − A| ÷ ΣA29.5%Defined at zero demand, weighted by the volume that actually costs money, and the same number whether you compute it on units or on cost. Says nothing about direction: a forecast 29.5% too high and one 29.5% too low score identically.
BiasΣ(F − A) ÷ ΣA, sign retained−12.1% (−5.6 units per week)The systematic component, and the only one of the four that is directly correctable. Says nothing about dispersion: a forecast that alternates +50 and −50 has zero bias and is useless.
Tracking signalΣ(F − A) ÷ MAD, where MAD = Σ|F − A| ÷ n−4.93 at week 12Detects drift rather than magnitude — it answers whether the errors have stopped cancelling. Needs a control limit set in advance, and needs a reset rule, or a long series makes it insensitive.

All four computed on the same twelve illustrative weeks above. The values differ by eight points on identical data, which is the point: an accuracy figure quoted without naming its metric is not comparable to any other accuracy figure.

MAPE’s asymmetry is the property most people are surprised by, and it is structural rather than incidental. The actual sits in the denominator, so an over-forecast against small demand has no ceiling: forecast 45 against an actual of 1 reads as 4,400% error. An under-forecast is bounded at 100%, because the largest possible under-forecast is zero. In week 8 of the example the forecast missed by 27 units below — 27 divided by 72, or 37.5%. Reverse the same 27-unit miss so the forecast is 72 against an actual of 45 and the identical error reads 60%. A metric that scores the same miss two ways depending on its direction will, applied over a season, steer a team toward chronic under-forecasting, because under-forecasting is the cheaper mistake in the report and the more expensive one in the warehouse.

WAPE has none of those problems and one of its own: it is direction-blind. A forecast running 29.5% high and one running 29.5% low are indistinguishable, and the two demand opposite responses — one ends in markdown, the other in an emergency chase against a sell-through signal read too late. That is the whole argument for publishing bias next to it rather than choosing between them.

Twelve weeks, one style-color, all four metrics

One style-color from a five style-color class, twelve trading weeks, in units. The forecast was frozen at the buy commitment and never revised. Week 4 is a genuine zero-demand week — the unit was on hand and sold nothing. Weeks 9 and 10 are part-week stockouts, where on-hand ran out and the recorded actual is capped by supply rather than by demand. These inputs were chosen because they divide cleanly.

Hypothetical twelve-week series for one style-color in units, showing forecast, actual, signed error, absolute error, the per-period percentage error MAPE uses, and whether the style-color was in stock all week.
WeekForecastActualError (F − A)Abs error|e| ÷ AIn stock
W13026+4415.38%Full
W23034−4411.76%Full
W33528+7725.00%Full
W4350+3535undefinedFull
W54052−121223.08%Full
W64058−181831.03%Full
W74566−212131.82%Full
W84572−272737.50%Full
W95048+224.17%Part
W105051−111.96%Part
W114562−171727.42%Full
W124055−151527.27%Full
Total485552−6716310 of 12

Illustrative figures, chosen because they divide cleanly. Not benchmarks, and not drawn from any brand. Forecast totals 485 units, actual totals 552, the signed error sums to −67 and the absolute error sums to 163. The final column records whether the style-color was on hand for the whole week — it is an input to the measurement, not an output of it.

Take the four metrics in turn. WAPE is 163 divided by 552, or 29.5%. Nothing has to be excluded and nothing has to be special-cased; the zero week contributes its 35 units of error to the numerator and its zero to the denominator, which is exactly right, because forecasting 35 units into a week that sold none is a real 35-unit miss.

MAPE cannot be computed at all. Week 4 requires dividing 35 by zero. The universal workaround is to delete that period, which leaves eleven weeks whose percentage errors sum to 236.39 points, so MAPE reads 21.49% — call it 21.5%. That figure is roughly eight points better than WAPE on identical data, and the reason is worth spelling out, because it is not a rounding artefact. Deleting week 4 removes 35 of the 163 units of absolute error from the calculation entirely: a fifth of the season’s total miss vanishes because of a division rule, and it is disproportionately over-forecast error that vanishes, since a zero-demand week can only ever be over-forecast. Compute WAPE on the same surviving eleven weeks and it reads 128 over 552, or 23.2%. The gap between 23.2% and 29.5% is the deleted week; the gap between 21.5% and 23.2% is the unweighted average letting the small weeks pull.

Bias is the signed sum over the actual: −67 divided by 552, or −12.1%, which is −5.6 units per week. That single figure carries information neither absolute metric contains — the forecast is systematically low. Look down the error column and the story is plainer still. The first four weeks run positive, then every remaining week except one runs negative. This is not noise around a correct level; it is a forecast that was right at the start of the run and progressively too low as demand built, which is a shape problem in the curve rather than a level problem in the total.

None of the three, however, would tell you when that drift began. That is what the tracking signal is for.

The tracking signal catches what the accuracy report cannot

A tracking signal is cumulative error divided by mean absolute deviation, recalculated every period. Both terms are running figures: the numerator keeps the sign so errors of opposite sign cancel, the denominator does not. An unbiased forecast produces a numerator that hovers near zero however large the individual misses, so the signal stays flat. A forecast developing a lean produces a numerator that walks steadily away from zero while the denominator grows only with the average size of a miss.

Running tracking signal across the twelve illustrative weeks: period error, cumulative error, cumulative absolute error, running mean absolute deviation, and the resulting signal.
WeekErrorCum. errorCum. |error|MADTracking signal
W1+4+444.00+1.00
W2−4084.000.00
W3+7+7155.00+1.40
W4+35+425012.50+3.36
W5−12+306212.40+2.42
W6−18+128013.33+0.90
W7−21−910114.43−0.62
W8−27−3612816.00−2.25
W9+2−3413014.44−2.35
W10−1−3513113.10−2.67
W11−17−5214813.45−3.86
W12−15−6716313.58−4.93

Same illustrative series. MAD is the running mean of the cumulative absolute error, so at week 12 it is 163 divided by 12, or 13.58, and the signal is −67 divided by 13.58, or −4.93.

Read the signal column downward and it tells a story the season-end WAPE erases completely. It runs positive through week 5 — the forecast was too high early — peaks at +3.36 in week 4 on the back of the zero-demand week, crosses zero in week 7, and then walks steadily negative to −4.93. The forecast did not have one problem. It had two opposite problems in sequence, and the season-end bias of −12.1% is the net of them, which understates both. A team acting only on the bias figure would shade the whole curve upward and make the early weeks worse.

The value of a limit is that it converts a chart into a decision. One common convention sets the limit at plus or minus four, beyond which the forecast is treated as out of control and re-fitted rather than nudged. On this series the signal is inside that band until week 11 at −3.86 and breaches it at week 12 at −4.93 — one week before the run ends, which is too late to act on. That is not an argument against the tool; it is an argument for running it at a grain and cadence where a breach still leaves trading weeks. The exact limit matters far less than fixing it before the season starts, because a limit chosen after seeing the data is not a control limit, it is a rationalisation.

One more property is worth naming: the signal becomes less sensitive as the series lengthens, because a long history of cancelling errors builds a large denominator that a recent lean has to overcome. Long-running replenishment series need a reset rule — restart at the start of each season, or run the calculation on a rolling window — or a forecast that went wrong in week 40 of 52 will never breach anything.

Bias is correctable, error mostly is not

Error and bias look like two views of the same thing and behave like two different problems. Bias is the systematic component: the forecast sits above or below the truth on average, and moving the forecast by the size of the bias removes it. It is arithmetic, it is cheap, and it usually has a nameable cause — a return assumption applied on the wrong basis, a promotional lift left in the baseline, a size curve carried forward from a season that fit differently, a launch date that moved and a curve that did not.

Error is dispersion: the spread of individual misses around whatever the forecast’s central level is. Some of it is model quality and can be reduced. Most of it, in a seasonal assortment, is the irreducible variability of the demand itself — weather, a piece of unpaid coverage, a competitor’s markdown, a store closure — and no amount of effort moves it. A team that reports absolute error alone works hard on the half of the problem it cannot fix and never sees the half it can. That is the single most common way a planning organisation spends a year improving its forecasting process and finishes with the same numbers.

The practical test is cheap. Take any population with a large absolute error and check whether the signed errors sum to something close to zero. If they do, the forecast is unbiased and noisy, and the answer is not a better forecast — it is a plan that tolerates the noise, which means buffer stock, a chase reserve, or a later commitment date. Sizing that tolerance is what safety stock for seasonal assortments and receipt flow planning are for. If the signed errors do not sum near zero, you have a correction to make, and it is usually a bigger prize than anything the model will give you.

The corollary is that bias and error should never be combined into a single score. A composite that trades one against the other lets a team improve the headline by making the correctable failure worse in exchange for a marginal reduction in the irreducible one, which is precisely backwards.

Roll it up and the number improves on its own

The first of the three things that make an accuracy number meaningless is the level it was summed at. Extend the worked example to the full class — five style-colors over the same twelve weeks — and compute the same WAPE at three grains without touching a single forecast.

Hypothetical class-level dataset: forecast, actual, signed error and summed weekly absolute error for five style-colors.
Style-colorForecast (12 wks)Actual (12 wks)Error (F − A)Σ weekly |error|
SC-1485552−67163
SC-2420396+24108
SC-3610638−28142
SC-4260214+4696
SC-5375400−25111
Class total2,1502,200−50620
The same forecast and the same actuals measured at three aggregation levels, showing WAPE falling as the grain coarsens.
GrainCells gradedΣ absolute errorActual unitsWAPE
Style-color × week606202,20028.2%
Style-color × season51902,2008.6%
Class × season1502,2002.3%

Same five style-colors, same twelve weeks, same forecast, same actuals. Nothing was re-forecast between rows — only the grain the error is summed at. Illustrative figures throughout.

Follow the numerator down the table. At style-color by week, sixty separate absolute errors sum to 620 units. Collapse the weeks and each style-color’s twelve errors net against each other before the absolute value is taken — SC-1’s 163 units of weekly error become a 67-unit season error — and the five season errors sum to 190. Collapse the style-colors too and the remaining five net down to a single 50-unit class error. The denominator never changes; it is 2,200 actual units at every level. So WAPE falls from 28.2% to 8.6% to 2.3%.

This is not an empirical tendency, it is the triangle inequality: the absolute value of a sum can never exceed the sum of the absolute values, with equality only when every error carries the same sign. Accuracy is therefore guaranteed to be non-decreasing as you aggregate, for any forecast, always. Which means a team that moves its reporting from SKU-week to class-month shows a large improvement having changed nothing at all, and a vendor comparing its class-level figure to your SKU-level figure is not making a comparison.

The bottom row deserves a second look. At class-season grain there is exactly one cell, so the absolute error and the signed error are the same number, and WAPE degenerates into bias: 50 over 2,200 is 2.3% either way. The fully aggregated accuracy figure is bias wearing a different name, and it is the figure most likely to reach a slide. It tells you the class was bought about right in total. It tells you nothing about whether the right style-colors were bought, which is the entire job.

The rule that follows is not to grade at the finest possible grain — at SKU-size-week most cells are zero and the measurement becomes noise. It is to grade at the grain the decision was made at. A buy commitment made at style-color depth should be graded at style-color; an allocation decision made at store-size should be graded there; a financial plan set at class-month should be graded at class-month. Grading above the decision grain flatters; grading below it measures a decision nobody made.

The other two ways the number becomes meaningless

Horizon is the second. A forecast made one week out and a forecast made six months out are different products with different achievable error, because the second one has to predict things the first can simply observe — the weather, the competitive set, whether the range landed on time, what the first four weeks of sell-through said. Reporting them under one heading averages two distributions that have nothing to do with each other, and the average moves whenever the mix of horizons in the report moves. State the horizon, and state it as the lag at which the decision was committed rather than the lag at which the comparison is convenient.

The third is the one that quietly voids most accuracy reporting in practice: measuring against a plan that has since been reforecast. A forecast that has been revised in light of what happened cannot be graded against what happened — the revision has already absorbed the answer. A team that reforecasts weekly and compares the final version to actuals is measuring how fast its plan converged on the present, and that measurement improves every time the cadence tightens. It is possible to report near-perfect accuracy this way while buying no better than the year before.

The fix is a freeze. Snapshot the forecast at the moment a decision was committed against it, store it immutably with its date, and grade that. Most plans need at least two frozen versions: the pre-season forecast the buy was placed against, and the in-season forecast the chase or markdown was placed against. They answer different questions. The first says whether the commitment was right; the second says whether the in-season read was right. A team that only holds the second never learns anything about its buying.

Freezing costs nothing and is skipped constantly, because in a spreadsheet the forecast column is the same column that gets updated. Once it has been overwritten the evidence is gone, and no amount of analysis afterwards recovers it.

Accuracy improves exactly when service fails

Every accuracy calculation compares a forecast of demand to a record of sales, and those are only the same thing while stock is available. When on-hand runs out, sales stop and demand does not. The recorded actual is censored — capped at what could be served — and the forecast is graded against a truncated number.

The direction of the distortion is the dangerous part. A forecast that was too low leads to a buy that was too small, which leads to a stockout, which caps the actual near the level that was bought — which is near the forecast. The under-forecast produces the stockout, and the stockout then hides the under-forecast. Weeks 9 and 10 of the worked example are exactly this: a forecast of 50 against recorded actuals of 48 and 51, contributing 2 and 1 units of absolute error against 99 units of volume. On the raw numbers they are the two most accurate weeks of the season. They are the two weeks the business lost money.

The same twelve-week series measured three ways: as recorded, filtered to in-stock weeks only, and with censored weeks reconstructed to an estimate of unconstrained demand.
TreatmentWeeks gradedActual unitsΣ absolute errorWAPEBias
As recorded1255216329.5%−12.1%
In-stock weeks only1045316035.3%−15.0%
Censored weeks reconstructed1261822536.4%−21.5%

The reconstruction assumes unconstrained demand of 80 and 85 units in weeks 9 and 10 against a 50-unit forecast — an illustrative estimate, labelled as one. Any reconstructed figure is an estimate and has to be reported as an estimate; the honest alternative is the middle row, which throws information away but invents nothing.

Work the middle row. Removing weeks 9 and 10 takes 99 units out of the denominator, leaving 453, and only 3 units out of the numerator, leaving 160. WAPE goes from 29.5% to 35.3%, and bias from −12.1% to −15.0%. Both numbers get worse, and both are more honest — the correction to a censored measurement always moves in the direction of a worse-looking result, which is why it is almost never applied voluntarily.

The bottom row goes further and reconstructs what would have sold. Estimating unconstrained demand at 80 and 85 units against the 50-unit forecast turns two near-perfect weeks into 30- and 35-unit misses, pushing WAPE to 36.4% and bias to −21.5%. Note which of the two moved most: the raw bias figure understated the systematic under-forecast by nearly half. That gap is the whole reason to do the correction. If the reported bias is −12.1% a planner nudges the curve; if it is −21.5% a planner rebuilds it.

Choosing between the two treatments is a question of how much you trust the reconstruction. Filtering to in-stock periods invents nothing but discards information, and it biases the remaining sample toward periods where demand was low enough not to break availability. Reconstruction keeps the periods but replaces an observation with an estimate, and every estimate needs its method stated — a rate extrapolated from the in-stock days of the same week, a comparable style running unconstrained, a pre-stockout run rate carried forward. Either is defensible. What is not defensible is leaving censored periods in unmarked, because that is the one option where the number gets better the worse you trade.

The practical prerequisite is availability history at the grain you are grading. If you know the forecast and the sales but not whether the item was on hand, this correction cannot be made at all, and the accuracy number is uninterpretable in the periods that matter most. That history has to be captured while it is happening; it cannot be reconstructed from a closing stock file at season end.

Eight steps, in order

The steps most often skipped are the first and the fifth — freezing the forecast, and handling censored periods. A measurement missing either one produces a number that improves for reasons unconnected to the forecast.

  1. 1

    Freeze the forecast and record the date it was frozen

    Copy the forecast to an immutable snapshot at the moment a decision was committed against it, and grade that snapshot. A forecast that has been revised since cannot be graded, because the revision has already absorbed the information you are testing it against.

  2. 2

    State the grain the error is summed at

    Decide whether you are grading style-color by week, style by month, or class by season, and name it. Error is non-increasing as you aggregate, so the same forecast produces different numbers at every level and an unqualified figure is a choice of level rather than a measurement.

  3. 3

    State the horizon being graded

    A one-week-ahead forecast and a six-month-ahead forecast are different products. Grade the horizon at which the decision was actually made — for a seasonal buy that is the commitment date, not the week before delivery.

  4. 4

    Use WAPE for magnitude and report bias beside it

    WAPE weights the error by the volume that costs money and stays defined when demand is zero. Bias keeps the sign. Neither substitutes for the other, so publish them together as a pair rather than choosing between them.

  5. 5

    Handle censored periods explicitly

    Identify every period where availability capped the actual, then either grade only in-stock periods or reconstruct unconstrained demand and label the reconstruction as an estimate. Both are defensible; leaving censored periods in unmarked is not.

  6. 6

    Segment before you average

    Split new lines from carry-over, promoted periods from non-promoted, and booked volume from forecast volume before computing anything. An average across two unrelated demand regimes describes neither and moves when the mix moves.

  7. 7

    Run a tracking signal with a limit set in advance

    Compute cumulative error over mean absolute deviation each period and compare it to a limit agreed before the season. The limit is a convention, so its exact value matters far less than the fact that it was fixed before anyone saw the result.

  8. 8

    Publish the spec alongside the number

    Every accuracy figure travels with its metric, grain, horizon, freeze date, availability rule and population. A figure without those six cannot be compared to last season, to another category, or to anyone else.

Six fields that travel with every number

An accuracy figure is only comparable to another figure measured the same way, so the specification is part of the result rather than documentation about it. Six fields are enough. Put them in the header of the report, not in an appendix.

The six fields a forecast accuracy report must state, with an example value and the reason each is required.
FieldExample valueWhy it is required
MetricWAPE, with bias reported beside itDetermines how the error is weighted and whether direction survives. MAPE, WAPE and a simple accuracy percentage give different answers on identical data.
GrainStyle-color × weekError is non-increasing as you aggregate, by arithmetic rather than by improvement. Without the grain the number is a choice of level, not a measurement.
HorizonTwelve weeks ahead, measured at the buy commitmentA one-week-ahead forecast and a six-month-ahead forecast are different products with different achievable error. Grade the horizon at which the decision was actually made.
Freeze dateForecast as at 2026-04-06, not revisedA forecast graded after it has been reforecast measures nothing. The frozen version is the one the buy was placed against.
Availability filterIn-stock weeks only; censored weeks listed separatelyStockouts cap the actual, so unfiltered accuracy improves exactly when service fails. State the rule and state how many periods it removed.
PopulationCarry-over core only; new lines and promoted weeks reported separatelyMixing intermittent and continuous items, or promoted and non-promoted weeks, produces an average of two unrelated distributions that describes neither.

Written out for the worked example, the whole spec is one line: WAPE and bias, style-color by week, twelve-week horizon, forecast frozen at the buy commitment, in-stock weeks only, carry-over core excluded. That line makes 35.3% and −15.0% reproducible by someone else and comparable to the same class next season. Without it, 29.5% is a number that cannot be checked, cannot be repeated, and cannot be argued with — which in practice means it gets quoted more confidently, not less.

Same arithmetic, different grain

The metrics do not change by vertical. What changes is the grain at which demand is continuous enough to measure, what censors it, and which populations have to be separated before anything is averaged. No accuracy figures appear below, because a credible one has to come from your own history measured to your own spec.

How the measurable grain and the dominant measurement problem differ across seven verticals.
VerticalGrain that can be measuredWhat distorts the number
ApparelStyle-color × week for volume; size share measured separately.The style-color-size-week grid is mostly empty cells, so MAPE is undefined across most of it and any average built on the non-empty cells is a selected sample. Short lifecycles mean there is often no in-season reforecast to grade, so the frozen forecast is the pre-season one and the horizon is the whole season.
FootwearStyle × week for volume; size run graded as share error.A style can be accurate in total and wrong in every size. Broken size runs censor demand at size level long before the style shows a stockout, so an availability filter applied at style level removes nothing and the size-level error stays hidden.
Accessories & bagsSeparate populations for carry-over core and seasonal colour.Evergreen core is continuous and low-intermittency; hero colourways are short and spiky. Averaged together the core carries the number and the fashion error disappears into it. Gifting concentration also means the horizon that matters ends at the buy commitment, not one week out.
Home & furnitureSKU × month, or collection × month, rather than SKU × week.Unit velocity per SKU is low enough that most SKU-weeks are zero and even WAPE is noisy at that grain. Long lead times mean the decision horizon is months, not weeks, and where a shortfall becomes a quoted lead-time extension rather than a lost sale the demand is captured, so an in-stock filter matters less than a lead-time flag.
OutdoorCategory × week, with the weather regime recorded alongside.A large part of the error is driven by conditions the forecast could not have known. Without recording those conditions next to the frozen forecast, a bad season and a bad forecast are indistinguishable, and the team re-tunes a method that was not the problem.
Health & beautyShade × month for established lines; new launches reported separately.Shade ranges create the same intermittency as size runs, and launch cadence means a large share of volume has no history at all. Promotional intensity is the second problem: measuring accuracy without recording whether a promotion ran conflates forecast error with a change to the promo plan.
Sporting goodsBooked and forecast volume graded as separate populations.Dealer prebooks are commitments, not predictions. Grading booked volume as forecast accuracy inflates the number by however much of the season is pre-sold, and the inflation grows every time prebook penetration rises — which reads as a forecasting improvement and is not one.

Three of those need spelling out, because they change what the measurement is rather than where it is taken. In footwear the demand lives in the size run, so a style can be accurate in total and wrong in every size, and the two errors need separate measurements. Volume error is a WAPE on style-week units; size error is better measured as the deviation between forecast and actual size shares, because a size that took, say, 14% of the run against a forecast 10% is a curve error whether the style sold 500 pairs or 5,000 — figures used here only to make the shape of the error visible. The availability point is sharper still: a broken size run censors demand at size level while the style still shows stock, so an in-stock filter applied at style level removes nothing and the censored size cells stay in the average. The filter has to run at the grain the shopper actually shops. Building the shares themselves is covered in how to calculate size curves.

Home and furniture inverts the intermittency problem. Per-SKU velocity is low enough that a SKU-week grid is almost entirely zeros, so MAPE is unusable and even WAPE is dominated by a handful of cells; SKU-month or collection-month is usually the first grain where the measurement means anything. Lead times then set the horizon: the decision being graded was committed months before delivery, so a one-week-ahead comparison grades a forecast nobody bought against. And censoring works differently — where a shortfall converts into a quoted lead-time extension rather than a lost sale, the order is still captured, so the demand record is intact and an in-stock filter is the wrong correction. A quoted-lead-time flag is the right one, because the thing that suppressed demand was the promise date, not the empty shelf.

Sporting goods carries the measurement error that most inflates a headline number. Where dealer prebooks account for a large share of the season, that volume is a commitment rather than a prediction, and including it in an accuracy calculation grades the order book instead of the forecast. The inflation scales with prebook penetration, so a commercial shift toward booked business reads on the report as a forecasting improvement. Split booked and forecast volume into separate populations before computing anything, and grade the forecast line on its own. Model-year changeovers add a second discipline, since they reset the horizon and end the comparable series; carrying a tracking signal across a changeover measures two different products as one.

The remaining four share one instruction: separate the populations first. In apparel that means the style-color-size-week grid is too sparse to average across, so volume is graded at style-color-week and the size distribution separately, and a short lifecycle means the frozen forecast is usually the pre-season one with the whole season as its horizon. In accessories and bags it means carry-over core and seasonal colour are different demand regimes — the core is continuous and predictable, the hero colourways short and spiky — and averaging them lets the core carry the number while the fashion error disappears into it.

In health and beauty it means three splits, not one: established shades against new launches, promoted periods against non-promoted, and shade-level against line-level. A shade range behaves like a size run for intermittency purposes, launch cadence means a large share of volume has no history to forecast from at all, and heavy promotional intensity makes much of the demand a forecast of the promo plan as much as of the customer — so an accuracy report that does not record whether a promotion ran will attribute a promo-calendar change to forecasting skill. Outdoor has the same requirement pointed at a different variable: record the weather regime alongside the frozen forecast, because without it a bad season and a bad forecast are indistinguishable, and the team spends the off-season re-tuning a method that was never the problem.

Eight ways an accuracy number lies

Grading a forecast that has been reforecast

The plan is updated weekly, and at season end the final version is compared to actuals. That number measures how fast the plan converged on what already happened, not how good the forecast was. It improves every time the reforecast cadence tightens, which is why teams that reforecast more often report better accuracy and buy no better than before.

Quoting an accuracy number without a grain

Aggregation does not improve a forecast; it cancels errors against each other. The relationship is the triangle inequality — the absolute value of a sum is never greater than the sum of absolute values — so accuracy is guaranteed to be non-decreasing as you roll up, with equality only when every error carries the same sign. A team that quietly moves from SKU-week to class-season reporting shows a large improvement having changed nothing.

Using MAPE at SKU-week grain in a fashion assortment

Most SKU-week cells are zero, MAPE is undefined at zero, and the universal workaround is to drop those cells. The dropped cells are not random: they are the periods where the forecast was highest relative to demand, so deleting them systematically removes over-forecast error and reports a better number on a worse forecast.

Chasing error while ignoring bias

Error is dispersion and is mostly irreducible below the noise floor of the demand itself. Bias is systematic and is correctable by adjusting the forecast. A team that reports only an absolute-error metric works on the half of the problem it cannot fix, and never sees the half it can.

Treating a stockout week as a normal observation

When on-hand runs out, recorded sales are capped by supply. The forecast is compared against a truncated actual, the gap closes, and measured accuracy improves precisely because service failed. Nothing in the accuracy report flags it, and the worse the availability the better the number looks.

Reporting a single accuracy percentage with no direction

Over-forecasting by twenty percent and under-forecasting by twenty percent score the same on any absolute-error metric, and they demand opposite responses: one ends in markdown, the other in lost sales and an emergency chase. A number that cannot distinguish those two cannot be acted on.

Grading booked volume as if it were forecast

Where part of the season is pre-sold — wholesale bookings, dealer prebooks, contracted replenishment — that volume is a commitment, not a prediction. Including it inflates measured accuracy in proportion to how much of the business is booked, and the inflation grows as booking penetration grows, so a mix shift reads as a forecasting improvement.

Comparing your number to a published accuracy figure

Two accuracy percentages are only comparable when the metric, grain, horizon, freeze rule, availability treatment and population all match, and published figures almost never state those. A quoted number that beats yours may be measuring class-month WAPE against a reforecast plan while yours measures frozen style-color-week error.

Why this drifts back to the default

Teams that understand all of this still end up quoting one unqualified accuracy percentage, and the reason is mechanical rather than analytical. The measurement needs four things held together for every cell being graded: the forecast as it stood on a specific past date, the actual, the availability state in that period, and the population tags that say whether the item was new, promoted or pre-booked. In a spreadsheet the first of those is overwritten by the next reforecast, the third lives in a system nobody exports weekly, and the fourth exists only in someone’s head.

So the measurement collapses to what can actually be computed from the surviving data: the latest plan against the latest actuals, at whatever level the file happens to be at, with no availability filter and no segmentation. That produces a number, and the number is always flattering, because every shortcut in that list moves it the same way. Nothing in the process flags it — the arithmetic is correct, the file reconciles, and the figure is comparable to nothing including last year’s version of itself.

When the forecast, the plan, the receipts and the availability history sit on one data model, freezing is a write rather than a discipline, the availability state is already attached to the period, and the population tags are attributes rather than tribal knowledge. The spec becomes a query definition instead of a monthly reconstruction — which is the difference between an accuracy number a team acts on and one it recites.

See the connected workflow in RetailNorthstar

Frequently asked questions

How do you measure forecast accuracy in retail?
Pick a metric, a grain and a horizon, freeze the forecast at the point a decision was committed against it, and compare that frozen version to actuals at the stated grain. For a merchandising team the working default is WAPE — total absolute error divided by total actual units or value — reported alongside bias, which keeps the sign. Then handle availability explicitly, because periods where stock ran out have a capped actual and will otherwise flatter the result. The measurement is only reproducible if all six of those choices travel with the number.
What is the difference between MAPE and WAPE?
MAPE averages the percentage error of each period independently, so every period counts equally regardless of size, and it is undefined whenever an actual is zero. WAPE divides total absolute error by total actual, so periods are weighted by their volume and a zero-demand period is handled without special treatment. The practical consequence is that MAPE lets a low-volume week with a large percentage miss dominate the average, while WAPE weights the error by the thing that costs money. MAPE is also asymmetric: over-forecasting has no ceiling because the actual sits in the denominator, whereas under-forecasting cannot exceed 100%.
What is forecast bias and how is it different from error?
Bias is the mean error with its sign retained — the sum of forecast minus actual, divided by total actual. Error, in the sense that MAPE and WAPE measure, is the average magnitude of the miss with the sign thrown away. They are different failures with different fixes. Bias is systematic: the forecast is consistently high or consistently low, which means it can be corrected by moving the forecast. Error is dispersion around the truth, and below the noise floor of the demand itself it is largely irreducible. A team that tracks error alone works on the part it cannot change and never sees the part it can.
Why does MAPE not work at SKU level?
Because MAPE divides by the actual, and at SKU-week grain in a fashion or seasonal assortment most actuals are zero. Division by zero makes the metric undefined, and the standard workaround — dropping the zero periods — is not neutral. Zero-demand periods are exactly the periods where the forecast was highest relative to demand, so removing them strips out over-forecast error and reports a better figure on an unchanged forecast. The periods that survive are also the low-volume, high-percentage ones, which then dominate an unweighted average. Use WAPE at that grain, or move the measurement to a level where demand is continuous.
What is a good forecast accuracy percentage?
There is no publishable number, and any figure quoted without its conditions is not comparable to yours. Four things set the achievable band before skill enters at all: demand intermittency, since a SKU that sells in most periods is far more predictable than one with mostly zero weeks; aggregation level, since error is arithmetically non-increasing as you roll up; horizon, since a one-week-ahead forecast and a six-month-ahead commitment are different problems; and promotional and event intensity, since demand driven by a promotion is a forecast of a plan as much as of a customer. Category structure sits underneath all four. The usable comparison is your own number against your own number, measured the same way, season over season — which is why the reporting spec matters more than the benchmark.
What is a tracking signal?
A tracking signal is cumulative forecast error divided by mean absolute deviation, recalculated each period. It answers a different question from an accuracy metric: not how big the misses are, but whether they have stopped cancelling. A forecast that is unbiased produces errors of both signs that sum toward zero, so the signal stays near zero however large the individual misses. When the signal drifts steadily in one direction and breaches a limit set in advance, the forecast has developed a systematic lean and needs re-fitting rather than more effort. The limit is a convention agreed before the season, and it must be fixed before anyone has seen the result.
Does a stockout make your forecast look better?
Yes, and that is the most dangerous property in the whole measurement. Recorded sales are capped by what was available, so when on-hand runs out the actual stops climbing toward the demand that existed. If the forecast was too low, the constrained actual lands close to it and the period scores as accurate. Measured accuracy therefore improves exactly when service fails, and it does so silently, because nothing in the accuracy report knows about availability. The correction is to grade only periods that were in stock throughout, or to reconstruct unconstrained demand for the censored periods and label the reconstruction as an estimate. Both make the reported number worse, which is the correct direction.

See how RetailNorthstar holds a frozen forecast, the actuals and the availability history on one data model, so accuracy is measured at the grain the decision was made at instead of the grain the file happens to be in.