The most common incrementality testing mistakes are stopping tests too early, using mismatched holdout groups, and treating non-significant results as evidence the campaign works. All three lead to bad budget decisions.
Incrementality testing is rigorous when done well and misleading when done carelessly. The stakes are high: a flawed test can lead you to cut a channel that was working or double down on one that was not. Understanding the common failure modes is as important as knowing the mechanics.
Mistake 1: stopping early when results look positive
Looking at results before the test is complete and then stopping because the numbers look good is called "peeking." It dramatically inflates your false positive rate. A test that shows 90 percent confidence on day 10 may show only 50 percent confidence on day 30 as the random noise averages out.
Decide on your test duration before you start. Commit to it. Do not check significance daily and stop at the first positive result you see.
Mistake 2: using mismatched holdout groups
A holdout group that was already converting at a higher rate before the test started will make your campaign look less effective than it is. A holdout group that was already lower-value will make your campaign look more effective than it is.
Always run a pre-test comparison. Look at at least four weeks of data before the test starts to confirm that your test and holdout groups were behaving similarly. If they were not, your matched groups were not a good match.
Mistake 3: contamination between groups
- Running campaigns that bleed across geographic boundaries, such as a digital campaign that targets a geo area without tight geo fencing
- Word-of-mouth from treated customers reaching holdout customers in nearby areas
- Promotional emails or app notifications that reach users across both groups
- Retargeting campaigns that follow users from treated regions into holdout regions based on prior website visits
Mistake 4: testing during unusual periods
Running an incrementality test during a promotional period, a product launch, or a major external event will contaminate the results. The lift you measure may reflect the event rather than the advertising.
Schedule tests during normal trading periods. If your business is highly seasonal, make sure the test spans the same seasonal window in both the test and holdout groups equally.
Treating a directional but non-significant result as evidence of effectiveness is one of the most common and dangerous mistakes. A positive number that is not statistically significant could easily be noise. Making a large budget decision on it is no better than guessing.
Mistake 5: changing the campaign mid-test
If you change your creative, targeting, bids, or budget during a test, you cannot cleanly interpret the results. The lift you measure at the end of the test reflects a combination of the original campaign and whatever changes you made. Keep the campaign static for the duration of the test.
Mistake 6: running underpowered tests
An underpowered test will almost never reach statistical significance even when a real effect exists. This causes teams to incorrectly conclude that a channel has no incremental value, when in reality the test simply lacked the data to detect the effect.
Always do a power calculation before starting. If the numbers show you need 12 weeks to reach adequate power and your test window is 4 weeks, you have a problem to solve before the test launches.
How do I know if my test was contaminated?
Look for telltale signs in the data: unusually similar conversion rates in both groups, unexpected sales spikes in the holdout that track the treatment group, or campaign delivery data that shows impressions reaching holdout regions. A strong pre-test period comparison can also reveal whether groups were equally exposed to other activity before the test began.
Can I salvage a test that was stopped early?
You can report the results with honest caveats: the test was stopped before reaching the planned duration, and the results should be treated as directional rather than conclusive. Do not present the early results as a validated finding. If the decision is important, run a full follow-up test.
My holdout group was small to reduce revenue risk. Did that compromise the test?
Possibly. A very small holdout group, such as 5 percent, reduces revenue risk but also reduces statistical power. If the test was underpowered as a result, the results may not be reliable. Balance revenue risk against the cost of an inconclusive test before choosing your holdout size.
