A spare part sells nothing for several days, then one customer buys six units. A conventional daily forecast dashboard can make that pattern look like repeated planning failure. A model that predicts zero every day may even look attractive on some error measures while leaving the business unable to serve the occasional order that matters.
EcomToolkit’s approach is to choose the decision horizon before choosing the forecast score. This guide focuses on intermittent demand for slow-moving ecommerce SKUs. The examples are hypothetical and the suggested workflow is an editorial framework. It is not a universal inventory policy or evidence that one forecasting method outperforms all alternatives.
Table of Contents
- Distinguish zero demand from missing opportunity
- See why percentage errors break
- Use WAPE with its denominator visible
- Match the horizon to replenishment
- Compare simple baselines and sparse demand methods
- Bring the result to the buying meeting
- EcomToolkit point of view
Distinguish zero demand from missing opportunity
Begin with the reason for a zero. A product that was available all day but attracted no purchases is different from a product that was unavailable, unpublished, blocked from a market, or missing from the reporting feed. Treating those conditions identically trains the forecast on a mixture of demand and operating failures.
Build a complete calendar at the SKU and location grain you actually replenish. Missing rows should not silently become zero sales. Record availability and eligibility alongside units so an analyst can inspect the observation process. When true demand is censored by stockouts, observed sales alone cannot reveal exactly how many units would have sold.
Also separate new products from established sparse products. A new SKU has limited history rather than necessarily intermittent long-run demand. An obsolete item may be approaching permanent inactivity. Ask merchandising whether the item is still part of the active assortment before interpreting every long gap as normal seasonality.

See why percentage errors break
Mean absolute percentage error divides each absolute forecast error by that period’s actual value. A zero actual value makes that term undefined. Removing every zero day changes the evaluation question: the remaining score describes only days on which demand occurred and ignores forecasts made during the gaps.
The forecast accuracy chapter in Forecasting: Principles and Practice explains this limitation and distinguishes test-set forecast errors from fitted residuals. For sparse products, retain the zero days and use measures whose interpretation remains meaningful for the question being asked.
Consider the hypothetical five-day actual series below. Forecast A predicts zero every day; forecast B predicts one unit every day. Neither is a recommended production model. The exercise shows how a seemingly simple score can reward an outcome that a replenishment team would not automatically prefer.
| Day | Actual units | Forecast A | Forecast B |
|---|---|---|---|
| 1 | 0 | 0 | 1 |
| 2 | 0 | 0 | 1 |
| 3 | 0 | 0 | 1 |
| 4 | 0 | 0 | 1 |
| 5 | 5 | 0 | 1 |
Forecast A has total absolute error of five and MAE of one unit. Forecast B has total absolute error of eight and MAE of 1.6 units. Yet B predicts the correct five-day total of five units, while A predicts no demand over the horizon. That does not automatically make B the better stocking policy; it shows why daily error and horizon requirements must be distinguished.
Use WAPE with its denominator visible
Weighted absolute percentage error is the sum of absolute errors divided by the sum of absolute actual values. In the example, A has WAPE of 100% and B has WAPE of 160%. Aggregating the denominator avoids division by each individual zero, but the result is still undefined when all actual observations in the evaluation window are zero.
Rob Hyndman’s discussion of WAPE explains additional limitations, including its statistical assumptions and problems for some changing series. Do not treat the ability to calculate a percentage as proof that it is the best measure for a sparse catalog.
Always show the number of evaluated periods, total actual units, and the count of all-zero windows. An unavailable result should remain unavailable. Replacing it with zero would make an unevaluable forecast look perfect. Replacing it with an arbitrary large number introduces an undocumented penalty that can dominate a portfolio ranking.
Portfolio aggregation deserves a separate label. Summing errors and demand across all SKUs weights the result toward higher-volume items. Averaging SKU-level percentages gives each evaluable SKU equal influence and excludes undefined cases unless you handle them separately. These are different statistics, so never compare them under the same dashboard name.
Match the horizon to replenishment
If a supplier needs two weeks to deliver, the daily forecast is only part of the planning problem. Evaluate cumulative demand over the lead-time horizon and consider the review interval used to place orders. A timing miss inside the horizon can matter less than a major error in the total, depending on available stock and service requirements.
Retain both horizon error and daily behavior. Aggregation can conceal a large early order that causes a stockout before replenishment arrives. An inventory simulation should represent the actual lead time, order constraints, and starting stock, rather than assuming that a correct period total guarantees uninterrupted service.
| Evaluation view | What it helps answer | Limitation to retain |
|---|---|---|
| Daily MAE in units | Typical absolute daily error | May favor conservative sparse forecasts |
| Lead-time total error | Quantity needed before replenishment | Hides within-horizon timing |
| Signed forecast minus actual | Systematic over or under prediction | Positive and negative errors cancel |
| Service outcome simulation | Whether demand could be fulfilled | Depends on inventory policy assumptions |
| Excess stock exposure | Cost of ordering too much | Requires holding and obsolescence assumptions |
For signed error, state the convention in the column heading. Here, positive forecast minus actual means overprediction. A different team may define error as actual minus forecast. The sign is not self-explanatory, and mixing conventions can reverse the apparent action in a buying review.

Compare simple baselines and sparse demand methods
A sensible evaluation includes transparent baselines before adding complexity. Depending on the series, these might include a historical mean, a seasonal rule supported by enough history, and a method intended for intermittent demand. Compare forecasts made with information available at the time, not models refitted using future observations.
The count series chapter describes Croston’s method, which separately smooths nonzero demand quantities and the intervals between them. It also notes limitations, including bias and the absence of a direct statistical model for standard prediction intervals. Treat the method as a candidate, not an automatic answer to every sparse series.
A fractional expected demand can be meaningful even when orders contain whole units. An expectation of 0.4 units per day does not claim that a shopper buys four-tenths of a spare part. Translate that estimate into an ordering decision through the replenishment policy, including pack sizes and minimum order quantities.
Use rolling evaluation cutoffs so a single unusual period does not decide the entire comparison. Avoid random train-test splits that allow later demand to inform earlier forecasts. Keep the same observation windows and product eligibility rules across candidates, and document where insufficient history prevents a fair comparison.
Bring the result to the buying meeting
Select a small group of active sparse SKUs with different unit costs and lead times. Review their time series, availability gaps, forecast alternatives, and simulated stock outcomes together. A rare expensive component may deserve a different stocking decision from a cheap accessory even when their percentage errors match.
Ask the buyer to challenge the assumptions explicitly: supplier reliability, substitution options, customer willingness to wait, pack sizes, and expected retirement date. Record those decisions beside the forecast version. The model provides an estimate; the operating policy expresses the business’s tradeoff between stock exposure and service.
Use the unit-of-measure guide before converting predictions into cases or cartons. Then connect the review to the broader demand and stock-risk framework without replacing the sparse-SKU evidence with a single catalog average.
EcomToolkit point of view
Zero sales are part of the signal, but only after the business establishes that a purchase was possible. A useful forecast scorecard keeps that distinction, the replenishment horizon, and the cost of being wrong visible. If your buying dashboard rewards forecasts that never order anything, request an EcomToolkit audit to review the metric and inventory decision together.