Back to the archive
Ecommerce Analyses

Conversion Improved Everywhere. Why Did the Store Total Fall?

Understand Simpson's paradox in ecommerce reporting and compare platform conversion statistics using segment rates, traffic weights, and worked tables.

An ecommerce operator reviewing performance metrics on a laptop.

A store’s mobile conversion rate rises. Its desktop conversion rate rises too. Yet the overall conversion rate falls. That result can be mathematically correct: more traffic has moved into the segment with the lower underlying rate. Before declaring a redesign unsuccessful or blaming an ecommerce platform, inspect the composition of the visitors behind the total.

EcomToolkit’s approach is to keep the observed business result and the like-for-like comparison side by side. Neither should erase the other. This guide explains Simpson’s paradox through a fully worked ecommerce example, then applies the same reasoning to platform comparisons and weekly reporting. All numerical examples are invented teaching data, not merchant results or published platform benchmarks.

Analysts reviewing ecommerce reporting and segment comparisons

Table of Contents

Start with a consistent conversion definition

For this example, conversion means a session with at least one completed purchase divided by eligible sessions. A session with two orders still counts once in the numerator. Order count per session is a different measure. Declare the reporting time zone, eligibility rules, bot treatment, attribution window, and source system before comparing anything.

The broader principle is documented in Mixpanel’s explanation of Simpson’s paradox: relationships in aggregated data can differ from relationships within subgroups. The calculation below is original and deliberately simple so a team can reproduce it with a spreadsheet.

Device is only one possible grouping. New versus returning visitors, country, channel, customer type, and product category can also change the composition of a store’s audience. Choose groups because they have a plausible relationship to both the outcome and the comparison. Searching hundreds of arbitrary segments until one tells a favorable story is not a reliable analytical method.

Calculate a reversal from the raw counts

Suppose Period A has mostly desktop sessions, while Period B has mostly mobile sessions. The mobile rate improves from 2% to 3%, and the desktop rate improves from 5% to 6%. Each period contains 10,000 eligible sessions, which keeps the arithmetic transparent without being necessary for the effect.

Period and deviceEligible sessionsPurchasing sessionsConversion rateShare of period traffic
A: mobile2,000402.0%20%
A: desktop8,0004005.0%80%
A: total10,0004404.4%100%
B: mobile8,0002403.0%80%
B: desktop2,0001206.0%20%
B: total10,0003603.6%100%

The overall rate decreases by 0.8 percentage points, from 4.4% to 3.6%. Relative to Period A, that is approximately an 18.2% decrease. Neither statement contradicts the one-percentage-point improvement within each device group. The share of sessions arriving through the lower-converting device rose sharply.

A simple average of 2% and 5% would produce 3.5%, but that is not Period A’s observed conversion rate. The overall rate must reflect the session counts. Always aggregate purchasing sessions and eligible sessions first, then divide. Do not average row percentages unless the weighting scheme is intentional and explicitly labelled.

Hold the traffic mix constant

To ask how conversion would compare under a common composition, choose a reference mix. In this example, use Period A’s 20% mobile and 80% desktop weighting for both periods. Multiply each segment’s conversion rate by its reference weight and add the results.

Period A remains 20% multiplied by 2%, plus 80% multiplied by 5%, which equals 4.4%. Period B becomes 20% multiplied by 3%, plus 80% multiplied by 6%, which equals 5.4%. At that fixed mix, the later period improves by one percentage point.

Reported viewPeriod APeriod BChange
Observed conversion4.4%3.6%−0.8 percentage points
Conversion at A’s device mix4.4%5.4%+1.0 percentage point
Mobile session share20%80%+60 percentage points
Actual purchasing sessions440360−80 sessions

The standardized number is a comparison device, not an alternative sales total. Period B did not actually produce 540 purchasing sessions. It produced 360. The fixed-mix calculation answers a conditional question about rates under selected weights; it cannot establish how customers would behave if acquisition were changed to create that audience.

Keep the reference mix stable during the reporting series. If you change the reference, restate the comparison or start a clearly labelled series. A moving reference can make a dashboard hard to audit and obscure the very composition change you intended to explain.

Separate performance change from composition change

A compact decomposition makes the example useful in a management meeting. First measure the rate change using the original weights. That contribution is plus one percentage point. Next measure the weight change using the later segment rates: mobile gains 60 percentage points of weight at a 3% rate, while desktop loses the same weight at a 6% rate. The composition contribution is minus 1.8 percentage points.

Together, plus 1.0 and minus 1.8 equal the observed decline of 0.8 percentage points. This is an exact arithmetic decomposition for the stated method. Another valid method can allocate the interaction between changing rates and changing weights differently, so document the convention instead of presenting one allocation as uniquely correct.

The business discussion now becomes more precise. Did a campaign bring additional early-stage mobile visitors? Did desktop acquisition fall? Did inventory or geography change alongside device mix? Those are hypotheses requiring evidence. The table identifies where to investigate; it does not explain why the composition changed or whether the acquisition decision was profitable.

Team discussing differences between aggregate and segmented results

Apply the method to platform statistics

A published claim that one ecommerce platform converts better than another needs more than two headline percentages. Ask whether the samples have comparable countries, merchant sizes, product categories, returning-customer shares, traffic sources, and devices. Also ask whether a conversion means an order, a purchasing session, or a checkout completion. Identical labels do not guarantee identical measures.

Even a device-adjusted comparison is incomplete if one platform’s sample contains mostly established subscription brands and the other contains new stores acquiring cold traffic. Standardization addresses the variables included in the calculation. It does not remove unmeasured differences or turn observational platform statistics into a randomized experiment.

For a migration review, maintain the same reporting definitions before and after launch and record simultaneous changes in promotions, consent collection, analytics implementation, and catalog availability. Compare operational outcomes as well as conversion. A platform decision also involves cost, maintainability, and the team’s ability to run the store. The platform selection guide provides that wider decision context.

Do not take the illustrative figures above and relabel Period A as Shopify or Period B as WooCommerce. They establish an arithmetic possibility, not an advantage for any real platform. A responsible comparison needs independently described samples and verified source definitions before a numerical ranking becomes useful.

Build a report that survives scrutiny

Put the observed total, segment counts, segment rates, fixed reference weights, and standardized result in the same report. Include an unknown-device row if the source contains one. Dropping unknown observations only in one period changes the denominator and can produce an artificial improvement.

Reconcile the segment counts to the total. Check that each eligible session belongs to exactly one mutually exclusive device group. When using multiple dimensions, watch for sparse cells and unstable estimates. More detail is not automatically more reliable; a tiny group with one purchase can swing dramatically between periods.

Keep sample uncertainty separate from the composition calculation. The arithmetic can be exact for observed counts while the underlying rates remain uncertain. For an experiment, use its planned statistical analysis and check allocation quality. Do not adjust for a behavior caused by the treatment itself merely because it makes the result look cleaner.

The real-user monitoring scorecard offers a related lesson for speed reporting: changing audience composition can change the headline. Conversion rates can be standardized with explicit weights; percentiles require their own distribution-aware treatment. Do not average subgroup percentiles and assume the result is the population percentile.

Questions about interpretation

Does the paradox mean the total is wrong? No. The total describes the observed population correctly. It is the interpretation that may be incomplete.

Should we always use the earlier period as the reference? No. A pooled or strategically defined population may be useful. Choose it in advance, explain why, and keep it stable.

Does better performance in both segments prove the redesign worked? No. Seasonality, campaign quality, inventory, and other changes can improve both groups. Causal attribution needs a stronger design.

What should leadership do next? Investigate the source and economics of the mix change while preserving the evidence of within-segment improvement. The right action might concern acquisition rather than the storefront.

EcomToolkit point of view

A useful ecommerce report explains how the business result arose. Show what happened, show how the audience changed, and make the comparison assumptions visible. That lets a team respond to the actual problem instead of optimizing for a misleading headline. For help reviewing your reporting definitions, request an ecommerce audit.

Related partner guides, playbooks, and templates.

Related ecommerce guides.

Free Shopify Audit

Get a free Shopify audit focused on the fixes that can move revenue.

Share the store URL, the blockers, and what needs attention most. EcomToolkit will review UX, CRO, merchandising, speed, and retention opportunities before replying.

What you get

A senior review with the priority issues most likely to improve performance.

Best for

Brands planning a redesign, migration, CRO sprint, or retention cleanup.

Reply route

Every request is routed to info@ecomtoolkit.net.

We use these details to review your store and reply with the next best steps.