Back to the archive
Ecommerce Analytics

Did the Campaign Create Sales? Ecommerce Geo Experiments and Synthetic Controls

Design and review ecommerce geo experiments with credible counterfactuals, pre-period fit, placebo checks, incremental revenue, and margin-aware iROAS.

An ecommerce operator reviewing performance metrics on a laptop.

Platform-reported revenue answers which orders received credit under an attribution rule. It does not establish how many orders would have disappeared without the campaign. That missing counterfactual is why a profitable-looking channel can remain incrementally weak.

Geo experiments create a practical route to causal measurement when user-level randomization is unavailable. Selected regions receive a treatment while other regions help estimate what would have happened otherwise. Synthetic-control methods can combine control regions into a closer counterfactual, but the model does not rescue a poorly designed test.

Table of Contents

Define the intervention and outcome

Write the treatment precisely: channel, audience rules, creative set, spend range, start time, end time, and geographic delivery controls. “Run more paid social” is not reproducible. A useful protocol might specify a four-week increase in prospecting spend in ten treatment regions with existing brand and retention activity held within agreed bounds.

Select the primary outcome before looking at results. Net product revenue, contribution margin, new-customer orders, or qualified leads can each be valid, but they answer different questions. Gross platform revenue can overstate value when cancellations, returns, discounts, taxes, and fulfillment costs differ across regions.

Define event time consistently. Order-created date supports acquisition analysis; fulfilled or returned states mature later. If the outcome will be restated after refunds, specify the maturation window in advance. Otherwise the apparent lift can shrink after the decision has already been celebrated.

Team planning an ecommerce measurement experiment

Choose regions before choosing a model

Regions need enough outcome volume, measurable media delivery, and limited spillover. Tiny geographies create noisy daily data. Adjacent areas with heavy commuting can contaminate one another. National promotions, press coverage, store openings, stock constraints, and weather can create treatment-control differences unrelated to media.

Start with a long enough pre-period to understand seasonality and co-movement. Match or stratify regions using outcome level, trend, volatility, customer mix, and other pre-treatment variables that are available consistently. Do not match on a post-treatment value.

Google Research describes randomized paired geo experiments for incremental return on ad spend and notes the challenges of small numbers of heterogeneous geographies. Its Trimmed Match research is a useful primary reference for robust paired-geo analysis.

Keep a design table before launch. It should show region assignment, pre-period revenue, spend, population or customer base, excluded dates, and known operational events. This prevents convenient redesign after the outcome is visible.

Build and test the counterfactual

A synthetic control assigns weights to eligible control regions so their weighted pre-period outcome resembles the treatment region or group. The post-period difference between observed treatment and predicted counterfactual becomes the estimated effect under the model assumptions.

Good pre-period fit is necessary but not sufficient. A model can overfit noise, rely excessively on one control, or fail when relationships change. Inspect the time series, residuals, weight concentration, and stability across alternate pre-periods. Keep the model specification fixed before reading the treatment-period effect.

Review questionUseful evidenceWarning sign
Does the control track before treatment?Low and stable pre-period errorVisible trend divergence
Are weights credible?Several relevant controls contributeOne fragile region dominates
Is the effect timed plausibly?Gap begins after treatmentGap begins before launch
Is the result stable?Similar under declared sensitivity checksSign flips across minor choices

Google’s time-based regression work describes predicting a counterfactual market response for geo experiments. See Estimating Ad Effectiveness using Geo Experiments for the underlying causal-measurement framing.

Calculate incremental economics

Assume an illustrative test produced 1.26 million in observed treatment-region net revenue and a synthetic counterfactual of 1.18 million. Estimated incremental revenue is 80,000. If incremental ad spend was 32,000, incremental ROAS is 2.5.

That is not yet profit. Suppose contribution before incremental media is 42% of net revenue. Incremental contribution is 33,600. After 32,000 of incremental media, the estimated contribution after media is 1,600. The campaign can look strong on revenue while providing a very narrow economic margin.

Illustrative componentAmountInterpretation
Observed treatment revenue1,260,000Actual defined outcome
Estimated counterfactual1,180,000Modelled no-treatment outcome
Incremental revenue80,000Observed minus counterfactual
Incremental spend32,000Treatment spend above baseline
Incremental ROAS2.50Revenue lift divided by spend
Contribution after media1,60080,000 × 42% − 32,000

All values above are hypothetical. Real decisions need uncertainty around the effect and a contribution model appropriate to the business. If the plausible effect range includes zero or negative margin, report that decision risk directly.

Use placebo checks and uncertainty

Placebo tests ask whether the method finds effects where no treatment occurred. Reassign pseudo-treatment dates within the pre-period or apply the design to untreated regions. If the method frequently produces gaps as large as the reported lift, the apparent campaign result is not distinctive.

Check for pre-trends. A treatment-control gap that starts before media changes is evidence against the intended causal story. Also inspect individual daily residuals rather than only the cumulative chart; a single launch day or outage can dominate the total.

Do not treat a point estimate as certainty. Publish an interval or a design-appropriate inference result, plus the assumptions used to obtain it. Distinguish statistical uncertainty from operational uncertainty such as an unrecorded local promotion.

Pre-register the primary analysis and a limited sensitivity set. Trying many outcomes, windows, exclusions, and model variants until one becomes positive creates a multiple-testing problem. The false discovery rate guide explains why repeated searching raises false-positive risk.

Analysts reviewing regional ecommerce results

Control contamination and interference

Geo experiments assume that treatment in one unit does not materially change outcomes in control units. Ecommerce breaks that assumption through cross-border shopping, commuting, shared media markets, word of mouth, and customers shipping outside their home region.

Define geography using the unit that governs exposure, not whichever address is easiest to query. Media delivery may use inferred location while revenue uses shipping address. Measure mismatch rates where possible and state the expected direction of dilution.

National inventory is another interference path. Treatment demand can consume stock that would have sold in control regions. The measured effect then includes reallocation, not only new demand. Monitor stockouts, delivery promises, and allocation rules during the test.

Freeze overlapping campaigns or at least log them. A creator campaign, email push, marketplace promotion, or price change that targets regions differently can invalidate the simple comparison. Operational discipline is part of experimental validity.

Turn the result into a budget rule

Decide what evidence changes spend before the test begins. The rule can require positive incremental contribution, a minimum lower-bound iROAS, or a result strong enough to justify a larger replication. Avoid a binary “worked/failed” label when the test is underpowered.

Use the result within the tested range. A positive effect from adding 32,000 does not prove that adding 320,000 will preserve the same return. Saturation, auction prices, reach, and creative fatigue change with scale.

Store the full design packet with the decision: treatment definition, assignments, exclusions, model version, outcome extract, placebo results, effect interval, and business action. Future tests should learn from it rather than starting from a screenshot of a platform dashboard.

The EcomToolkit view

Geo experiments are strongest when design leads and modelling follows. A synthetic control should make the counterfactual more credible, not make weak assumptions less visible. Tie the effect to contribution margin and a pre-agreed decision rule.

If attribution reports claim success but cannot estimate the no-campaign outcome, request an EcomToolkit audit to build an incrementality measurement plan.

Related partner guides, playbooks, and templates.

Related ecommerce guides.

Free Shopify Audit

Get a free Shopify audit focused on the fixes that can move revenue.

Share the store URL, the blockers, and what needs attention most. EcomToolkit will review UX, CRO, merchandising, speed, and retention opportunities before replying.

What you get

A senior review with the priority issues most likely to improve performance.

Best for

Brands planning a redesign, migration, CRO sprint, or retention cleanup.

Reply route

Every request is routed to info@ecomtoolkit.net.

We use these details to review your store and reply with the next best steps.