Platform-reported revenue answers which orders received credit under an attribution rule. It does not establish how many orders would have disappeared without the campaign. That missing counterfactual is why a profitable-looking channel can remain incrementally weak.
Geo experiments create a practical route to causal measurement when user-level randomization is unavailable. Selected regions receive a treatment while other regions help estimate what would have happened otherwise. Synthetic-control methods can combine control regions into a closer counterfactual, but the model does not rescue a poorly designed test.
Table of Contents
- Define the intervention and outcome
- Choose regions before choosing a model
- Build and test the counterfactual
- Calculate incremental economics
- Use placebo checks and uncertainty
- Control contamination and interference
- Turn the result into a budget rule
- The EcomToolkit view
Define the intervention and outcome
Write the treatment precisely: channel, audience rules, creative set, spend range, start time, end time, and geographic delivery controls. “Run more paid social” is not reproducible. A useful protocol might specify a four-week increase in prospecting spend in ten treatment regions with existing brand and retention activity held within agreed bounds.
Select the primary outcome before looking at results. Net product revenue, contribution margin, new-customer orders, or qualified leads can each be valid, but they answer different questions. Gross platform revenue can overstate value when cancellations, returns, discounts, taxes, and fulfillment costs differ across regions.
Define event time consistently. Order-created date supports acquisition analysis; fulfilled or returned states mature later. If the outcome will be restated after refunds, specify the maturation window in advance. Otherwise the apparent lift can shrink after the decision has already been celebrated.

Choose regions before choosing a model
Regions need enough outcome volume, measurable media delivery, and limited spillover. Tiny geographies create noisy daily data. Adjacent areas with heavy commuting can contaminate one another. National promotions, press coverage, store openings, stock constraints, and weather can create treatment-control differences unrelated to media.
Start with a long enough pre-period to understand seasonality and co-movement. Match or stratify regions using outcome level, trend, volatility, customer mix, and other pre-treatment variables that are available consistently. Do not match on a post-treatment value.
Google Research describes randomized paired geo experiments for incremental return on ad spend and notes the challenges of small numbers of heterogeneous geographies. Its Trimmed Match research is a useful primary reference for robust paired-geo analysis.
Keep a design table before launch. It should show region assignment, pre-period revenue, spend, population or customer base, excluded dates, and known operational events. This prevents convenient redesign after the outcome is visible.
Build and test the counterfactual
A synthetic control assigns weights to eligible control regions so their weighted pre-period outcome resembles the treatment region or group. The post-period difference between observed treatment and predicted counterfactual becomes the estimated effect under the model assumptions.
Good pre-period fit is necessary but not sufficient. A model can overfit noise, rely excessively on one control, or fail when relationships change. Inspect the time series, residuals, weight concentration, and stability across alternate pre-periods. Keep the model specification fixed before reading the treatment-period effect.
| Review question | Useful evidence | Warning sign |
|---|---|---|
| Does the control track before treatment? | Low and stable pre-period error | Visible trend divergence |
| Are weights credible? | Several relevant controls contribute | One fragile region dominates |
| Is the effect timed plausibly? | Gap begins after treatment | Gap begins before launch |
| Is the result stable? | Similar under declared sensitivity checks | Sign flips across minor choices |
Google’s time-based regression work describes predicting a counterfactual market response for geo experiments. See Estimating Ad Effectiveness using Geo Experiments for the underlying causal-measurement framing.
Calculate incremental economics
Assume an illustrative test produced 1.26 million in observed treatment-region net revenue and a synthetic counterfactual of 1.18 million. Estimated incremental revenue is 80,000. If incremental ad spend was 32,000, incremental ROAS is 2.5.
That is not yet profit. Suppose contribution before incremental media is 42% of net revenue. Incremental contribution is 33,600. After 32,000 of incremental media, the estimated contribution after media is 1,600. The campaign can look strong on revenue while providing a very narrow economic margin.
| Illustrative component | Amount | Interpretation |
|---|---|---|
| Observed treatment revenue | 1,260,000 | Actual defined outcome |
| Estimated counterfactual | 1,180,000 | Modelled no-treatment outcome |
| Incremental revenue | 80,000 | Observed minus counterfactual |
| Incremental spend | 32,000 | Treatment spend above baseline |
| Incremental ROAS | 2.50 | Revenue lift divided by spend |
| Contribution after media | 1,600 | 80,000 × 42% − 32,000 |
All values above are hypothetical. Real decisions need uncertainty around the effect and a contribution model appropriate to the business. If the plausible effect range includes zero or negative margin, report that decision risk directly.
Use placebo checks and uncertainty
Placebo tests ask whether the method finds effects where no treatment occurred. Reassign pseudo-treatment dates within the pre-period or apply the design to untreated regions. If the method frequently produces gaps as large as the reported lift, the apparent campaign result is not distinctive.
Check for pre-trends. A treatment-control gap that starts before media changes is evidence against the intended causal story. Also inspect individual daily residuals rather than only the cumulative chart; a single launch day or outage can dominate the total.
Do not treat a point estimate as certainty. Publish an interval or a design-appropriate inference result, plus the assumptions used to obtain it. Distinguish statistical uncertainty from operational uncertainty such as an unrecorded local promotion.
Pre-register the primary analysis and a limited sensitivity set. Trying many outcomes, windows, exclusions, and model variants until one becomes positive creates a multiple-testing problem. The false discovery rate guide explains why repeated searching raises false-positive risk.

Control contamination and interference
Geo experiments assume that treatment in one unit does not materially change outcomes in control units. Ecommerce breaks that assumption through cross-border shopping, commuting, shared media markets, word of mouth, and customers shipping outside their home region.
Define geography using the unit that governs exposure, not whichever address is easiest to query. Media delivery may use inferred location while revenue uses shipping address. Measure mismatch rates where possible and state the expected direction of dilution.
National inventory is another interference path. Treatment demand can consume stock that would have sold in control regions. The measured effect then includes reallocation, not only new demand. Monitor stockouts, delivery promises, and allocation rules during the test.
Freeze overlapping campaigns or at least log them. A creator campaign, email push, marketplace promotion, or price change that targets regions differently can invalidate the simple comparison. Operational discipline is part of experimental validity.
Turn the result into a budget rule
Decide what evidence changes spend before the test begins. The rule can require positive incremental contribution, a minimum lower-bound iROAS, or a result strong enough to justify a larger replication. Avoid a binary “worked/failed” label when the test is underpowered.
Use the result within the tested range. A positive effect from adding 32,000 does not prove that adding 320,000 will preserve the same return. Saturation, auction prices, reach, and creative fatigue change with scale.
Store the full design packet with the decision: treatment definition, assignments, exclusions, model version, outcome extract, placebo results, effect interval, and business action. Future tests should learn from it rather than starting from a screenshot of a platform dashboard.
The EcomToolkit view
Geo experiments are strongest when design leads and modelling follows. A synthetic control should make the counterfactual more credible, not make weak assumptions less visible. Tie the effect to contribution margin and a pre-agreed decision rule.
If attribution reports claim success but cannot estimate the no-campaign outcome, request an EcomToolkit audit to build an incrementality measurement plan.