The key outputs from an incrementality test are: lift percentage, confidence interval, and incremental revenue. All three matter. A lift number without a confidence interval tells you very little about whether you can trust the result.
Many marketers receive incrementality reports and focus entirely on the headline lift number. That is a mistake. A 15 percent lift that could plausibly range from 2 percent to 28 percent is very different from a 15 percent lift that ranges from 12 percent to 18 percent. The confidence interval tells you how much to trust the number.
Lift percentage
Lift is the percentage increase in the outcome metric in the test group relative to the holdout group. A lift of 12 percent means the people or regions exposed to your advertising converted 12 percent more than those who were not.
Lift is a relative measure. A 12 percent lift on a base conversion rate of 2 percent produces very different absolute revenue from a 12 percent lift on a base of 10 percent. Always translate lift into absolute revenue to understand the business value.
Confidence intervals and statistical significance
A confidence interval gives you a range within which the true lift likely falls. A 95 percent confidence interval means that if you ran this test 100 times, the true value would fall within the interval in approximately 95 of those runs.
Statistical significance tells you the probability that a result this large, or larger, would occur by chance if the true lift were zero. A p-value below 0.05 is the standard threshold for calling a result significant, though this is a convention, not a hard rule.
What non-significant results mean
A non-significant result does not mean the campaign had no effect. It means you cannot reliably distinguish a real effect from random variation with the data you collected. This could mean the campaign truly does nothing, or it could mean the test was underpowered.
Before concluding that your campaign is ineffective, check whether the test had adequate statistical power for the lift size you observed. An underpowered test cannot produce a significant result even when a real effect exists.
The most common mistake in reading results is treating a non-significant positive lift as evidence the campaign works. Without statistical significance, you cannot distinguish lift from noise. Do not optimise based on a directional number that is not significant.
Incremental revenue and iROAS
Once you have a reliable lift estimate, translate it into incremental revenue by multiplying the lift percentage by the baseline revenue from the holdout group, scaled to represent the full market. Then divide by the spend in the test group to get your incremental ROAS.
- Incremental revenue = Holdout group baseline revenue x Lift %
- Incremental ROAS = Incremental revenue / Campaign spend
- Compare iROAS to your target ROAS to decide whether the campaign justifies its budget
How to act on the results
If lift is statistically significant and iROAS is above your threshold, the campaign is working and scaling is justified. If lift is significant but iROAS is below threshold, the campaign is incremental but not efficient enough. If the result is not significant, do not change budgets based on the test. Instead, redesign and retest with more power.
What should I do if my confidence interval crosses zero?
An interval that crosses zero means you cannot rule out that the true lift is zero or even negative. This is a non-significant result. Do not present the midpoint of the interval as your lift estimate and base decisions on it. You need a larger or longer test to reach a reliable conclusion.
Can I compare lift numbers across different channels?
With care, yes. Lift percentages are comparable across channels as long as the holdout methodology was the same. Be cautious comparing platform-run audience holdouts to independently run geo holdouts: the measurement methods differ and the numbers may not be directly comparable.
My results show negative lift. What went wrong?
Negative lift means the holdout group outperformed the test group. This can happen due to a poorly matched holdout group that was already higher-value, contamination in the data, or very rarely, genuinely negative advertising effects such as ad fatigue. Review the pre-test period to check whether both groups were truly comparable before the test started.
