Incrementality testing: work out what your test can detect first.
A geo holdout with ordinary week-to-week variance cannot see a 10% effect no matter how long you run it. Teams run one anyway, get a null, and cut a channel that was working. The arithmetic that prevents this takes ten minutes.
Attribution answers which touchpoint gets credit for a conversion. Incrementality answers whether the conversion would have happened anyway, which is the question that actually decides budget. No amount of attribution modeling gets you there, because all of it is arithmetic over the same observational data: a channel can claim thousands of conversions while creating almost none of them, and every model you run will keep agreeing with itself. Causality needs a control group.
The standard advice stops there, hands you three test designs, and tells you to run them for four to six weeks. The missing step is the one that determines whether the exercise is worth doing at all: working out what size of effect your data is physically capable of detecting before you spend six weeks and a chunk of revenue finding out. Skip it and you will eventually run a test that could never have found anything, read the null as evidence of nothing, and cut a channel that was working. That failure mode is common, expensive, and entirely preventable with ten minutes of arithmetic.
The arithmetic that decides whether to run the test
The minimum detectable effect of a two-arm test is driven by how noisy the metric already is and how many periods you observe. For eighty percent power at ninety-five percent confidence the approximation is tight enough to plan with:
MDE = 2.8 * CV * sqrt(2 / weeks)
CV = weekly standard deviation of revenue
in the test regions, as a share of
their mean
AT 15% WEEKLY VARIATION
4 weeks -> MDE 29.7%
6 weeks -> MDE 24.2%
12 weeks -> MDE 17.1%
26 weeks -> MDE 11.6%
AT 25% WEEKLY VARIATION
6 weeks -> MDE 40.4%
12 weeks -> MDE 28.6%
AT 8% WEEKLY VARIATION
6 weeks -> MDE 12.9%
12 weeks -> MDE 9.1%Read the middle column and the implications land immediately. A six-week geo holdout on a business with fifteen percent weekly revenue variation, entirely ordinary, can only detect a swing of roughly a quarter. Ask that test whether a channel is contributing a ten percent lift and it will return a null every time, regardless of the truth. And because nulls get reported as "no measurable impact", the test does not merely fail to answer the question. It answers it wrongly, in a consistent direction, with the authority of an experiment.
This is why underpowered incrementality testing is worse than none. Its errors are not random. They all point toward "nothing is incremental", because that is what insufficient power looks like. A team running badly sized tests will systematically defund its working channels and feel rigorous doing it.
Variance reduction is worth four times what duration is
Look at how the two inputs enter the formula. Variance enters linearly; weeks enter under a square root. Halving your coefficient of variation does the same work as quadrupling the length of the test. Every hour of design effort spent reducing noise is worth roughly four hours of patience, which reverses where most teams put their attention.
- Match control regions on pre-period correlation, not on size — Two regions of similar revenue that move differently week to week are a worse pair than a large and a small region whose weekly patterns track each other closely. Correlation is the property that cancels noise; size is not.
- Aggregate small regions into arms rather than picking a couple of big ones — Pooling ten cities into each arm averages away idiosyncratic local noise that two cities cannot.
- Exclude promotional weeks from the window, deliberately and in advance — A sale inside one arm and not the other is variance you chose to accept.
- Measure the ratio between arms rather than each arm’s own trend — A shared shock (a heatwave, a national news cycle, a supply problem) moves both arms together and cancels in the ratio while wrecking a comparison against historical baseline.
- Use a longer pre-period than test period — The pre-period is free variance reduction: it is what lets you establish the arms were comparable, and it costs no test time at all.
This also tells you which tests to run first. Do not start with the channel you are most curious about. Start with the one where the expected effect is largest relative to your MDE, because that is the test most likely to return an answer you can actually read, and a first readable answer is what earns the budget for the rest of the program.
Design one: the brand-search shutoff
Start here, and the reason is the arithmetic above rather than convenience. Pause brand search in a subset of geographies and watch what happens to total brand-driven conversions there. Organic listings usually catch a large share of the demand the ads were claiming, and the share that does not come back is the true incremental contribution of brand advertising. Crucially, the expected effect is enormous (often the majority of claimed conversions turn out to be non-incremental), which puts it far above the MDE of even a short, noisy test. It is the one design that reliably produces a legible result on the first attempt.
Fear of this test is diagnostic. If pausing brand ads for four weeks in two cities feels unthinkable, the account is running on faith in a number nobody has ever verified. The discomfort is the finding, before the test even starts.
Its real cost is competitive rather than statistical: going dark on your own name invites conquesting, and in categories where competitors bid aggressively on brand terms the lost revenue is not purely a measurement expense. Pick regions where that exposure is tolerable, and treat the result as a floor on brand ads’ value rather than a verdict. The reporting counterpart to this test (separating the two intents so the blended figure stops lying to you year-round) is in brand vs non-brand: the measurement error.
Design two: the geo holdout
Turn a channel off in a set of regions and leave it running everywhere else. If the channel is incremental, the dark regions should fall relative to the live ones; if sales hold up, the channel was harvesting demand rather than creating it. This is the general-purpose design and also the one most exposed to the power problem, so compute the MDE before committing.
It has one further limitation that survives any amount of statistical care, and it is worth stating plainly because it is almost never mentioned. Turning a channel entirely off measures the value of the whole channel, which is not the decision you usually face. You rarely want to know whether to delete Meta; you want to know what the next twenty percent of Meta budget is worth. Response to spend is non-linear and strongly diminishing, so the average effect of total removal tells you very little about the marginal effect at your current level.
Testing the margin directly (a thirty percent budget reduction in the holdout rather than a shutoff) is the honest design for that question, and it has a much smaller expected effect, which means much worse power. That tension is real and does not resolve. The practical answer is to test shutoffs to establish which channels matter at all, and to accept that marginal allocation stays a judgment informed by evidence rather than a measured quantity.
Design three: the audience holdout
For retargeting and CRM audiences, hold out a random slice and never show them the ads, then compare purchase rates across the same window. Randomization at the person level removes the geographic matching problem entirely, which makes this the cleanest of the three designs when the platform genuinely enforces the split.
That caveat carries weight. Held-out users can often still be reached by your other campaigns, so the contrast is between "targeted by this campaign" and "possibly reached anyway", which biases the measured effect toward zero. Check whether exclusions are actually being honored across every campaign in the account before believing a small result. Retargeting is the perennial over-claimer here. It specializes in reaching people who were already coming back, then charging you for their return, so a genuinely null retargeting holdout is usually a real finding rather than a power problem.
Reading results honestly
Pre-register the hypothesis, the metric, the revenue source it comes from, and the decision rule before launching. Change nothing else about the tested channel mid-flight. Then read the result against the MDE rather than against zero: an observed two percent effect from a test with a twenty-four percent MDE is not evidence of a two percent effect, and it is not evidence of no effect either. It is evidence that the effect, if any, is smaller than your instrument can see. Reporting that as "no impact" is the specific mistake that gets working channels cut.
A genuine null (an effect confidently below a threshold that matters) is a finding worth as much as a strong positive, because both move the next pound. The distinction between the two kinds of null is the entire discipline, and it is the part that gets skipped.
Where these tests cannot help
- Geo tests break on national campaigns and on spillover — People travel, ship to other addresses, and see national media. Where the treatment leaks across the boundary, the arms contaminate and the effect shrinks toward zero for reasons that have nothing to do with the channel.
- Cross-channel interference is not a nuisance, it is the norm — Pausing Meta in a region changes Google’s performance there. Whatever you measure is the effect of that channel in the presence of everything else you happened to be running, which is why results expire when the mix changes.
- Incrementality drifts — A result is a snapshot of a market, a creative set and a competitive landscape. Retest the significant channels once or twice a year, and treat any result older than that as a prior rather than a fact.
- Below a certain revenue base the arithmetic simply refuses — If your MDE at twelve weeks is still north of thirty percent, no test design will rescue it, and the honest move is to say so rather than to run something and report the number it produces.
- None of this fixes which revenue figure you are testing on — A geo test run against platform-reported conversions inherits every attribution problem it was meant to escape. Test against commerce truth. See MER vs ROAS for why the source choice dominates.
- The formula here is a planning approximation, not a statistical treatment — It is deliberately the version a marketing team can run in a spreadsheet, and it is accurate enough to tell you whether a test is worth starting. If a genuine seven-figure allocation rests on the result, the extra rigor of a proper design is cheap by comparison.
A brand that runs three or four well-sized tests a year makes structurally better budget decisions than one with the most sophisticated attribution dashboard money can buy, because it is the only one of the two that has ever observed its own counterfactual. But well-sized is doing real work in that sentence. Attribution allocates credit, testing establishes truth, and a test that cannot see the effect it was built to find establishes nothing while looking exactly like proof.
How long should an incrementality test run?
Long enough for the minimum detectable effect to fall below the effect you expect, which is a different question from a fixed number of weeks. Compute MDE as 2.8 times your weekly coefficient of variation times the square root of two over the number of weeks. At fifteen percent weekly variation, six weeks detects about a twenty-four percent swing and twelve weeks about seventeen percent. If the effect you are looking for is smaller than that, extending the test is an inefficient fix compared with reducing variance.
What is a minimum detectable effect and why does it matter?
It is the smallest true effect your test has a reasonable chance of finding, given how noisy the metric is and how long you observe it. It matters because a test cannot return "inconclusive" on its own. It returns a small number that reads like evidence of no effect. Knowing the MDE in advance is what lets you distinguish "this channel contributes little" from "this test could never have seen it", and those two conclusions lead to opposite budget decisions.
Which incrementality test should I run first?
A brand-search shutoff in a subset of regions. Not because it is easiest, but because the expected effect is unusually large (often the majority of conversions brand campaigns claim turn out to be non-incremental), which places it comfortably above the detection threshold of even a short, noisy test. It is the design most likely to produce a readable answer on the first attempt, and a readable answer is what earns the program its next test.
Can a geo holdout tell me how much budget to move?
Not directly, and this is the limitation least often acknowledged. Switching a channel fully off measures the value of the entire channel, while the decision you usually face is what the next increment of spend is worth. Because response to spend diminishes, the average effect of removal says little about the marginal effect at your current level. Testing a partial budget reduction addresses the right question but has a far smaller expected effect and therefore much weaker power.
Do you need a data science team to run these tests?
No. The three designs (brand-search shutoff, geo holdout, audience holdout) are all executable by a competent marketing team, and the planning arithmetic fits in a spreadsheet. What you do need is the discipline to compute the detectable effect before starting, to pre-register the decision rule, and to read a small result against that threshold rather than against zero. The statistics are not the hard part; the honesty about them is.
Written by Sam Nouri, founder, adsrunner. If this resonated and you want to apply it to your own account, you can book a strategy call or run a free audit.
How we research, source figures, and handle corrections: editorial policy.