Back to the archive
Ecommerce Analytics

Too Many Winners? Multiple Testing in Ecommerce Experiments

Control false discoveries when ecommerce experiments compare many metrics or segments, with worked thresholds and a practical plan for reporting reliable results.

An ecommerce operator reviewing performance metrics on a laptop.

A merchandising experiment looks neutral overall, so the team opens the country report, the device report, the category report, and several revenue metrics. Eventually one segment looks excellent. The result may be useful for a future hypothesis, but the process gave chance many opportunities to produce an attractive finding.

EcomToolkit’s view is that the search for a winner belongs in the analysis record. This article explains multiple testing through original hypothetical calculations and a proposed experiment review process. It does not claim that any specific store’s result is false or that a correction can rescue a poorly designed experiment.

Table of Contents

Count the questions you actually asked

A statistical test evaluates a particular hypothesis under assumptions. If a team runs many tests and only shows the favorable one, the audience cannot judge the selection process. The relevant family may include several variants, several outcomes, or many segment comparisons tied to the same decision.

Write the intended family before inspecting the results. For example, a planned comparison of five merchandising variants against a control is different from five unrelated experiments conducted for separate decisions. The choice of family should follow the scientific and commercial question, not whichever grouping produces a convenient threshold.

The NIST multiple comparisons guide explains why investigating several comparisons requires methods appropriate to that set of questions. A single comparison’s nominal error rate does not automatically describe the chance of a false claim across the full search.

Team reviewing a storefront workflow on a laptop

See how opportunities accumulate

Assume twenty independent tests, every null hypothesis true, and a five percent false-positive probability for each test. The probability of at least one false positive is 1 − 0.95^20, approximately 64.15%. Independence and true nulls are explicit assumptions in this worked example; real ecommerce metrics are often correlated.

The expected number of false positives in that setup is twenty times 0.05, or one. That expectation is not a promise of exactly one false result in each batch. Some batches produce none, others more. Do not reinterpret 64.15% as the probability that a particular observed winner is false.

Independent tests under true nullsPer-test thresholdChance of at least one false positive
10.055.00%
50.0522.62%
100.0540.13%
200.0564.15%

These values illustrate the selection problem, not the false-discovery rate of a real program. If tests share visitors or outcomes, the independence calculation is not exact. That is a reason to inspect the design, not a reason to ignore multiplicity altogether.

Choose the error criterion for the decision

Family-wise error control concerns the chance of making at least one false rejection within the defined family. False discovery rate control concerns the expected proportion of false discoveries among rejections, with that proportion taken as zero when there are no rejections. The two criteria answer different questions.

For a small set of consequential launch claims, a family-wise procedure may fit the review standard. For an exploratory screen that produces candidates for later confirmation, an FDR procedure may be useful. This is a proposed governance distinction, not a universal rule that every exploratory project should use the same method.

A Bonferroni threshold divides the family-level alpha by the number of tests. With five tests and alpha 0.05, the individual threshold is 0.01. The NIST Bonferroni reference explains the simultaneous-inference logic. Other procedures can provide different power properties while targeting the chosen error criterion.

Do not select the criterion after seeing which method approves your preferred variant. Record the decision first, along with the outcome definition, stopping rule, and treatment comparison. The methods are tools for answering a planned question, not interchangeable routes to a green dashboard cell.

Work through a Benjamini Hochberg screen

For a standard Benjamini–Hochberg procedure, sort the p-values and compare the value at rank i with i × q / m, where m is the test count and q the target FDR. Find the largest rank meeting its threshold and reject all hypotheses through that rank. This procedure requires valid p-values and appropriate dependence assumptions, commonly independence or specified positive dependence conditions.

Suppose the five sorted p-values are 0.003, 0.018, 0.041, 0.20, and 0.70, with q equal to 0.05. The second rank is the largest qualifying rank, so the first two hypotheses are selected. This is a toy calculation, not a complete analysis of experiment data.

RankRaw p-valueBH thresholdThreshold comparison
10.0030.010Meets threshold
20.0180.020Meets threshold
30.0410.030Does not meet threshold
40.2000.040Does not meet threshold
50.7000.050Does not meet threshold

A Bonferroni screen at the same numerical alpha would select only the first value in this example. Neither outcome establishes commercial value. The procedures target different error criteria, and an effect can be statistically detectable while too small or too expensive to implement.

The statsmodels multipletests documentation lists methods including fdr_bh, fdr_by, and holm. If using software, specify the method explicitly, preserve the original metric identifiers through sorting, and record the package version. Do not rely on a default method whose meaning the report never explains.

Colleagues discussing a performance analysis together

Keep exploratory segments honest

A post-hoc segment discovery can be valuable without being a confirmed launch result. Label it exploratory, retain the complete family of examined segments, and write the next test around a specific mechanism. A category may respond differently because of product fit, price, or availability rather than the interface change alone.

A difference between a significant result in one segment and a nonsignificant result in another is not itself evidence that the effects differ. Assess the interaction or treatment-effect contrast with an appropriate analysis. Otherwise, one small noisy group and one large precise group can create a misleading segmentation story.

Avoid turning every demographic or device slice into a separate shipping decision. Sparse segments create unstable estimates and maintenance complexity. Start with the customer problem and whether a distinct experience is justified, then decide what evidence would support that choice.

The archive’s confidence interval guide helps explain uncertainty around rates. Pair adjusted decision rules with effect sizes, intervals appropriate to the analysis, and eligible sample counts rather than reporting only corrected p-values.

Do not confuse multiplicity with other failures

Multiple-testing correction does not fix repeated unplanned peeking at a fixed-horizon test. It does not repair biased assignment, broken purchase events, or a metric defined after looking at treatment outcomes. Sequential monitoring needs a method designed for that monitoring plan.

Before correction, check assignment and data quality. The sample ratio mismatch guide covers a distinct diagnostic that can invalidate interpretation before any winner selection. A sophisticated correction applied to corrupted inputs remains an unreliable result.

Also inspect the analysis unit. Repeated visits by the same assigned shopper may require an analysis that respects the randomization unit. Treating every event as independent can produce invalid raw p-values, and adjusting those p-values afterward cannot restore the missing assumptions.

Publish a decision record people can review

A useful result record names the experiment, decision, primary outcome, test family, adjustment method, and stopping rule. Show all planned comparisons, including neutral and unfavorable outcomes. Explain any exclusions and distinguish prespecified work from exploration added during analysis.

Add commercial guardrails such as contribution, cancellation, and support burden where they are relevant to the change. An adjusted statistical result should inform the decision alongside implementation cost and customer impact. It should not automatically authorize a more complicated shopping experience.

For a follow-up test, state what new evidence would change the conclusion. Repeating a test until the same segment looks favorable is another selection process. A confirmation plan needs an independent opportunity to fail and a report that remains useful when it does.

EcomToolkit point of view

A credible experimentation program makes its search process visible. Fewer celebrated winners can mean better decisions when the apparent alternatives were selected from noise. If your reports contain many positive segments but few reproducible gains, request an EcomToolkit audit focused on test families, assignment, and the commercial meaning of the results.

Related partner guides, playbooks, and templates.

Related ecommerce guides.

Free Shopify Audit

Get a free Shopify audit focused on the fixes that can move revenue.

Share the store URL, the blockers, and what needs attention most. EcomToolkit will review UX, CRO, merchandising, speed, and retention opportunities before replying.

What you get

A senior review with the priority issues most likely to improve performance.

Best for

Brands planning a redesign, migration, CRO sprint, or retention cleanup.

Reply route

Every request is routed to info@ecomtoolkit.net.

We use these details to review your store and reply with the next best steps.