Incrementality testing is the only method that answers the question every other measurement tool dodges: would this conversion have happened anyway? Attribution tells you which touchpoint was nearest the sale. Platform reporting tells you what the platform is willing to claim. Incrementality tells you what the spend actually caused — and the gap between the three is usually larger than anyone on the account wants to say out loud.
What incrementality actually measures
Every incrementality test is built on the same idea: hold advertising back from a randomly chosen group, let it run for everyone else, and compare. The difference between the two is the causal effect. Everything else — the branded search that would have happened, the loyal customer who was going to reorder, the retargeted cart that was always going to convert — falls out of the number, because it happens on both sides.
That is the entire value proposition. It is also why incrementality results are so frequently unwelcome. A channel that reports a 4x return on platform-attributed conversions can turn out to be barely breaking even once the baseline is removed, and retargeting is the usual casualty.
The three designs that dominate
- Geo holdouts — turn spend off in a set of matched markets, keep it on elsewhere, and compare business outcomes. Channel-agnostic, works on any conversion you can measure by region, and unaffected by device-level signal loss.
- In-platform conversion lift — Meta and Google split an eligible audience into exposed and control groups, serving the control a ghost bid rather than the ad. Precise at the user level, but limited to that platform’s own view of the world.
- Synthetic control — build a weighted composite of untreated markets that tracks the treated market’s history, then measure the divergence after the change. Meta’s GeoLift package and Google’s CausalImpact both implement this, and it is the practical option when you cannot randomize cleanly.
Reading the result
The output is incremental lift, and from it two numbers you can act on: incremental CPA and iROAS — incremental revenue divided by the media cost in the test.
A worked example. A brand spends $500,000 in a quarter and the platform reports 4,000 conversions, a $125 CPA and a 2.4x return at a $300 average order value. A geo holdout puts the incremental figure at 2,300 conversions. That is a true incremental CPA of $217 and an iROAS of 1.4. The campaign is still profitable — but it is a different campaign from the one in the dashboard, and it should be budgeted like one.
What it costs, honestly
The holdout is not free: it is revenue you deliberately forgo to learn something. Common practice is a 10–20% holdout, large enough to produce a readable signal without turning the test into the quarter’s biggest line item. Tests also need scale — low-volume conversions produce confidence intervals so wide that the result cannot separate a good channel from a bad one.
And the answer decays. Creative changes, seasons turn, competitors move their spend. A single test is roughly a 90-day fact, which is why mature programs run a rolling schedule of six to twelve tests a year rather than treating one study as settled truth.
What incrementality cannot tell you
This is the limitation that matters most, and it is structural rather than fixable. Incrementality is measured on money already spent. It tells you that a channel returned 1.4x last quarter. It does not tell you which of the nine creatives in that flight carried the result, and it certainly does not tell you which of the six ideas you have not made yet is worth producing.
It also cannot run upstream. By the time a lift test reads out, the media is spent and the creative is live. If the weak asset in the rotation was the problem, incrementality will report a disappointing channel rather than a fixable ad — and the budget conversation that follows will be about cutting the channel.
The sequencing that works: test the concept before production, test the execution before launch, and reserve incrementality for the question it is uniquely good at — how much budget the channel deserves next quarter. Running it in place of pre-testing is how testing budgets get wasted.
Where it fits with MMM and attribution
The three are complements, not competitors, and each operates at a different altitude. Marketing mix modelling works on aggregate history and answers strategic allocation questions. Attribution runs daily and is useful for tactical pacing, provided nobody mistakes it for causality. Incrementality is the calibration layer between them: the experimental ground truth that tells you how much to trust the other two, and increasingly the input teams use to constrain their mix models.
None of them evaluates an ad. That is a separate job, and a cheaper one — creative testing answers it before the media is committed rather than a quarter after.
Frequently asked questions
What is incrementality testing?
An experiment that withholds advertising from a randomized control group in order to measure the conversions the advertising actually caused, as opposed to the conversions it was merely present for.
How is incrementality different from attribution?
Attribution assigns credit among the touchpoints a converter saw. Incrementality compares converters against people who saw nothing, so it measures causation rather than correlation. Attribution almost always reports a larger number.
How long does an incrementality test take?
Typically two to six weeks, depending on conversion volume and how large an effect you need to detect. Longer purchase cycles need longer tests, because a short window measures the timing of conversions rather than their existence.
What is a good iROAS?
There is no universal threshold — it depends on margin and on the payback period you can fund. The useful comparison is not against a benchmark but against the same channel’s reported ROAS, because the ratio between them tells you how much your dashboard is overstating.
Does incrementality testing replace creative testing?
No. It measures channels and budgets after the money is spent; creative testing chooses assets before it is. AdTest.AI scores an ad in about ninety seconds, so the creative decision is settled long before there is anything for a lift test to measure.