Measurement and attribution

Plan a geographic holdout test for paid media

A geographic holdout test compares business outcomes across regions receiving different advertising treatment. Define the causal question, validate historical regional data, assess detectable effect and spillover, and predefine analysis before launch. A regional split alone does not make an experiment reliable.

A geographic experiment can help a business estimate the effect of an advertising change when a suitable user-level experiment is unavailable or does not match the question. Regions receive different treatment, and their outcomes are compared under a planned design.

The hard part is not drawing a map with two colors. It is establishing that the treatment can be delivered, the outcome can be measured consistently, and the design can distinguish a decision-relevant effect from ordinary regional variation.

Define the treatment precisely

Write the advertising change the test will evaluate. Turning a channel off, introducing a new channel, and increasing spend on an existing channel estimate different effects.

Specify the campaigns, regions, budget behavior, dates, and any other settings that form the treatment. If the business also changes prices or promotions only in treatment regions, the experiment concerns a combined intervention unless the design explicitly separates them.

Use the attribution and incrementality guide to keep the causal question distinct from a request to reconcile platform credit.

Choose the business outcome and region definition

Select an outcome that the business can measure consistently across regions, such as completed orders or qualified opportunities. Define how an outcome is assigned geographically: customer residence, shipping destination, service location, or another appropriate field.

That assignment should fit the advertising treatment. A campaign targeting people physically present in a region may not align perfectly with orders classified by billing address. Investigate the mismatch before launch.

Keep the outcome definition stable during the test. A CRM qualification-policy change can make the measured series incomparable even if the regional ad settings remain correct.

Inspect the historical data

Build a region-by-time dataset with the outcome, relevant spend or exposure, and known business context. Look for missing periods, geography-code changes, one-time events, stockouts, and measurement releases.

Google's Meridian GeoX design guide uses pre-test data to evaluate candidate designs and explicitly notes that the tool does not automatically clean the input. Data preparation remains part of the experimenter's work.

Use historical periods that are relevant to expected test conditions. Removing inconvenient observations after seeing the result is different from a documented pre-test data-quality decision.

Assess whether the test can answer the question

Define the smallest effect that would change the business decision. Then evaluate the available regions, variability, duration, and budget against that requirement using an appropriate statistical design.

Minimum detectable effect is a design property, not the effect you hope to obtain. A test that can detect only a very large lift may be unhelpful when the business decision turns on a modest improvement.

The GeoX FAQ discusses tradeoffs involving the number of regions and restrictions on region assignment. Treat tool-specific recommendations as guidance for that method, not universal rules for every geographic experiment.

Review operational constraints before assignment

ConstraintPlanning question
Geographic deliveryCan the platform deliver the intended regional difference?
SpilloverCan travel, shared media, or national activity blur treatment?
Business capacityDo regions differ in inventory, service, or sales coverage?
Protected marketsAre exclusions necessary, and how do they affect the design?
Concurrent activityAre promotions or other channels changing by region?
MeasurementIs the same outcome collected consistently everywhere?

Do not choose the most convenient treatment regions and then assume they form a valid comparison. If business constraints restrict assignment, include those constraints in the design and assess the resulting limitations.

Prewrite the analysis and decision rule

Save the selected regions, treatment dates, primary outcome, analysis method, uncertainty level, exclusion rules, and planned readout date before launch. Record which secondary metrics are exploratory.

Define what would cause a pause for operational reasons, such as a service failure or incorrect geographic delivery. Those interventions may be necessary, but they can change the interpretation of the test.

Avoid repeatedly checking a result and stopping when it becomes favorable unless the analysis was designed for that sequential decision process. The stopping rule is part of the method, not an administrative detail.

Verify implementation during the test

Check that the intended regional spend or exposure difference actually occurred. Save settings and delivery evidence. A configured exclusion is weaker evidence than observed treatment separation.

Maintain a dated incident log for campaign edits, website failures, inventory changes, promotions, and geographic anomalies. Keep outcome monitoring separate from opportunistic redesign of the test.

If treatment delivery fails materially, report that fact. An inconclusive implementation can still reveal useful operational work, but it should not be presented as a clean estimate of advertising effectiveness.

Interpret the result within its scope

Report the estimated effect with uncertainty, treatment conditions, and any material deviations. “No statistically clear effect” is not automatically proof of zero effect. A wide interval may include both useful benefit and unacceptable downside.

An illustrative test might suggest positive lift but remain too imprecise to justify a large budget increase. The next decision could be a smaller bounded change, a redesigned experiment, or more observation, depending on the economics.

Use the spend scenario guide to translate a range of plausible effects into practical budget choices.

Preserve the experiment for later use

Store the pre-test design, input data version, implementation evidence, analysis, and final decision together. Future models or channel plans should receive the conditions and uncertainty, not just a single lift percentage.

The MMM readiness guide explains how experiments can become one input to broader measurement. A well-documented geographic test is valuable because another analyst can understand the question it answered and the limits on applying that answer elsewhere.

Use attribution and incrementality for different decisions

Separate advertising credit assignment from causal lift, and choose reporting, experiments, or modeling based on the budget question you need to answer.

Is your data ready for marketing mix modeling?

Assess marketing mix model readiness through consistent outcomes, media variation, time and geography, confounders, missing data, and calibration evidence.

Build three ad-spend scenarios for next month

Plan conservative, base, and expansion advertising scenarios with explicit response assumptions, contribution, capacity, cash requirements, and decision triggers.

Recover from a conversion-tracking outage

Diagnose and recover a conversion-tracking outage by tracing business events, containing unreliable automation, repairing the failing stage, and reconciling recovery.

Have a correction or a question about the workflow? Contact GaaS. Read our editorial standards for sourcing and example conventions.