A holdout test deliberately keeps a portion of your audience or market from seeing your advertising. By comparing their behaviour to people who did see ads, you can calculate how much of your sales were genuinely caused by the campaign.
Every advertising platform claims credit for conversions. Facebook says its ads drove the sale. Google says it did. The problem is they count using their own data, and both are right by their own rules. A holdout test bypasses that debate by measuring behaviour, not clicks.
The basic idea
You split your potential audience into two groups before the campaign starts. The test group receives the campaign as normal. The holdout group, typically 10 to 20 percent of your audience, is excluded from seeing the ads.
At the end of the campaign, you look at the conversion rate or revenue from each group. The gap between them, once you account for any baseline differences, is the incremental effect of your advertising.
Holdout vs control group: is there a difference?
In marketing, the terms are used interchangeably. A holdout group and a control group both refer to the unexposed segment of your audience. Some practitioners use "holdout" to describe geographic exclusions and "control" to describe audience-level exclusions, but there is no universal standard.
Types of holdout test
- Audience holdout: a random sample of users is excluded from seeing ads on a specific platform
- Geo holdout: entire regions are excluded from a campaign while matching regions continue as normal
- Channel holdout: all advertising for a single channel is paused in the holdout group while other channels run as normal
- Ghost ad test: users see a neutral placeholder ad rather than a real ad, ensuring equal exposure to ad inventory without the campaign message
What makes a holdout test reliable
The key requirement is that your holdout group must be comparable to your test group before the experiment starts. If the holdout group has systematically different purchase intent, the comparison will be misleading from the start.
Random assignment handles this at the audience level. For geo holdouts, you need to deliberately select matching regions based on historical sales patterns.
One important caveat: even well-designed holdout tests can be contaminated if people in the holdout group discuss your product with people in the test group. This "spillover" is a real problem for locally targeted campaigns in connected communities.
What holdout tests cannot tell you
A holdout test measures the total incremental effect of the campaign being tested. It does not tell you why the ads worked, which creative performed best, or how the effect varies across audience segments within the test group.
Holdout tests answer the question "did this campaign cause sales?" They do not answer "what should I do next?" You need additional analysis for that.
When to use a holdout test
Run a holdout test whenever you want to validate whether a channel is worth the budget you are putting into it. High-spend channels like paid social, TV, and display are the most common candidates because the stakes are highest and platform attribution is least trustworthy.
Holdout tests are also useful when you are about to make a significant budget change and want a reliable baseline before and after.
How large does my holdout group need to be?
The size depends on your conversion volume and the lift size you expect to detect. As a starting point, a 10 to 20 percent holdout is common. Smaller holdouts reduce revenue risk but require longer test periods or higher conversion volumes to reach statistical significance.
Will I lose revenue by running a holdout test?
Yes, potentially. The holdout group does not see your ads, so you may lose some sales from that group during the test. Think of this as the cost of getting reliable data. The insight is usually worth far more than the short-term revenue risk.
Can platforms run holdout tests for me?
Meta and Google both offer built-in holdout tests. They are convenient, but remember that these platforms run the test and report the results, which means there is an inherent conflict of interest. Platform-run tests tend to show higher lift than independently run tests. Use them as a starting point and verify with geo-based tests where possible.
