Back to the archive
Ecommerce Analytics

Should That Huge Order Stay? Outliers in Ecommerce Revenue Analysis

Handle ecommerce order-value outliers with a documented policy, compare raw and winsorized results, and preserve revenue truth while investigating unusual orders.

An ecommerce operator reviewing performance metrics on a laptop.

One unusually large order can move average order value enough to change a weekly trading discussion. The team may celebrate a merchandising improvement or remove the order as an outlier. Both responses are premature until someone establishes what the order represents and what the analysis is trying to measure.

EcomToolkit recommends preserving the raw revenue record and creating clearly labelled analytical views when needed. This guide offers an original decision process for unusual order values. All calculations are hypothetical. They illustrate how a method changes the question, not a benchmark for normal order size or a claim about a particular merchant.

Table of Contents

Decide whether the value is wrong or unusual

An unusually large order may be a valid wholesale purchase, a currency conversion error, a duplicated import, a test order, or a promotion-related bulk purchase. Those cases require different treatment. A statistical threshold cannot decide which business event occurred.

NIST’s outlier guidance distinguishes identifying unusual observations from investigating and accommodating them. It also cautions that some tests depend on distributional assumptions. In ecommerce, a highly skewed order-value distribution makes an automatic normal-distribution rule particularly worth questioning.

Start with source reconciliation: order identity, currency, payment state, line quantities, discounts, and known test flags. Do not edit source financial records merely to improve a dashboard. If a data pipeline is wrong, correct the pipeline with an auditable explanation and regenerate the analytical result.

Analysts examining an ecommerce workflow together

Define the business question first

A finance reconciliation asks what revenue the included orders actually produced under a defined accounting basis. A merchandising analysis may ask what a typical retail basket looks like. An experiment may ask about expected revenue per assigned customer. These are related questions with different denominators and sensitivity to large orders.

For typical basket size, report a median alongside the mean and distribution. For total revenue, a valid large order belongs in the total. For a causal experiment, changing extreme outcomes after seeing which group benefits can compromise interpretation even when the resulting chart looks more stable.

The archive’s metric grain guide is useful when the same customer places several orders. An outlier policy at order level is not equivalent to a policy at customer level. Define the analytical unit before deciding what counts as extreme.

Finding during reviewAppropriate next stepWhat to avoid
Duplicate importRepair and document deduplicationQuietly capping the duplicate value
Wrong currency conversionCorrect the conversion logicDeleting all foreign orders
Valid wholesale orderPreserve and consider segmentationCalling it an error because it is large
Known test transactionApply documented eligibility ruleInventing a rule after seeing results
Valid unusual retail orderRetain raw value and assess sensitivityTreating a threshold as proof of invalidity

Understand what winsorization changes

Winsorization replaces values beyond chosen cutoffs with boundary values. It retains observations rather than simply removing them, as described in NIST’s winsorization reference. The resulting mean describes transformed values, so it should not be labelled as the unmodified revenue mean.

Consider five hypothetical order values: 20, 30, 40, 50, and 360. Their total is 500 and their arithmetic mean is 100. Their median is 40. These are all correct summaries of the same small dataset, but they emphasize different features.

For illustration only, cap the largest value at a preselected boundary of 50. The transformed values total 190 and average 38. The store did not lose 310 in actual revenue; the analyst changed the values used for a particular summary. This distinction should be visible wherever the result is presented.

Show the raw and transformed views together

The following comparison uses that same five-order example. The cap is deliberately simple to make the arithmetic transparent. It is not a recommended percentile rule, and a sample of five orders is not sufficient to establish a production outlier policy.

Summary methodIncluded observationsResult
Raw totalAll five original values500
Raw arithmetic meanAll five original values100
MedianMiddle original value40
Mean after capping at 50Five transformed values38
Mean after removing 360Four remaining values35

Removing the high order changes the denominator as well as the total. Capping retains the denominator but changes the numerator. Neither result should replace the raw total in a revenue reconciliation. Their purpose, if justified, is to answer a separately named analytical question.

Include the number of affected observations and the amount transformed. A capped mean can appear stable while a growing share of revenue lies above the cap. That change may be commercially important, particularly when the store is attracting larger baskets or a new business customer segment.

Team discussing findings from a storefront analysis

Write the policy before comparing outcomes

A useful policy states the population, analytical unit, variable, cutoff method, and intended use. It explains whether cutoffs come from historical data, a business rule, or a prespecified statistical procedure. It also identifies who can change the rule and how old reports will be handled.

For an experiment, choose the analysis plan before examining treatment outcomes. If transformed metrics are used, retain a raw outcome view and explain the estimand: the precise quantity the analysis is intended to estimate. A capped revenue effect is not automatically the effect on total unmodified revenue.

Avoid choosing different cutoffs for groups merely because their distributions differ after treatment. The treatment itself may create larger orders. Group-specific transformation can obscure that change. A qualified analyst should choose uncertainty calculations and methods that match the design and the declared target quantity.

Segment only when the distinction is real

If wholesale and retail orders have different operating models, a separate segment can help. Define it using actual customer or order attributes rather than retroactively labelling every expensive basket wholesale. Otherwise the segment becomes another form of selective deletion.

Compare contribution, fulfillment cost, discounts, and return behavior where reliable data exist. A large order may be valuable, low margin, or operationally expensive. Order value alone cannot settle that question. Preserve the connection between the transaction and the downstream costs relevant to the decision.

Keep segment totals reconcilable to the complete population. A dashboard should show where the unusual orders went, not make them vanish between tabs. If a valid segment is excluded from a retail analysis, label that scope prominently in the title or methodological note.

Investigate changes in the tail

Track the count and share of unusually large orders over time using a stable definition. A sudden increase may reflect a promotion, new channel, data defect, or genuine change in buying behavior. The direction of the investigation depends on the supporting evidence.

Review example records with restricted access where customer information is involved. The editorial report does not need personal details to explain the finding. It needs the event type, reason for the classification, financial basis, and consequence for the metric.

A useful weekly handoff might state that three valid business orders account for a specified share of revenue and explain how the retail-only view differs. It should not claim that a smoother average is more truthful simply because it varies less from week to week.

Keep sensitivity analysis operational

Before acting, compare the decision under raw, segmented, and justified transformed views. If every view supports the same conclusion, the decision is less dependent on the unusual values. If the recommendation reverses, investigate the underlying transactions and communicate that dependence explicitly.

Do not run an unlimited menu of transformations until a preferred conclusion appears. Choose a small, documented sensitivity set that corresponds to plausible business interpretations. Save the query version and input period so another analyst can reproduce the comparison later.

The EcomToolkit view

An outlier is a prompt to understand a transaction. The strongest analysis preserves the economic truth and makes any transformation visible. A quieter dashboard is useful only when its metric still answers the question the team believes it is asking.

If a handful of orders repeatedly changes your trading decisions, request an EcomToolkit audit. Bring the metric definition and anonymized value distribution so the review can separate data errors, real customer segments, and analytical choices without erasing valid revenue.

Related partner guides, playbooks, and templates.

Related ecommerce guides.

Free Shopify Audit

Get a free Shopify audit focused on the fixes that can move revenue.

Share the store URL, the blockers, and what needs attention most. EcomToolkit will review UX, CRO, merchandising, speed, and retention opportunities before replying.

What you get

A senior review with the priority issues most likely to improve performance.

Best for

Brands planning a redesign, migration, CRO sprint, or retention cleanup.

Reply route

Every request is routed to info@ecomtoolkit.net.

We use these details to review your store and reply with the next best steps.