What Sample Size Do You Need for an Incrementality Test?

There is no single right answer to sample size. It depends on three things: how much lift you expect your ads to produce, how much natural variation exists in your sales data, and how confident you need to be in the result.

Running an incrementality test without thinking about sample size is like conducting a survey of five people and presenting the results as the national view. The test might point in the right direction, but you cannot trust it.

The three inputs to sample size

To estimate the sample you need, you must define three things before you start.

  • Minimum detectable effect (MDE): the smallest lift you care about. If your ads produce 5 percent incremental lift but your test can only detect effects above 20 percent, you will incorrectly conclude the campaign does nothing.
  • Statistical power: typically set at 80 percent. This means your test has an 80 percent chance of detecting a real effect if one exists.
  • Significance level: typically set at 95 percent confidence. This means you accept a 5 percent chance of seeing a positive result when none actually exists.

What drives up the required sample

High variance in your sales data means you need a larger sample to distinguish a real signal from normal fluctuation. Small expected lift means you also need more data, because the effect is closer to the noise floor.

This is why low-volume advertisers struggle with incrementality testing. If you generate fewer than a few thousand conversions per month, most standard tests will be underpowered unless the lift effect is very large.

Audience holdouts vs geo holdouts

For audience holdouts, sample size is measured in users. Most online experiments with conversion outcomes need tens of thousands of users per group to detect a 5 to 10 percent lift reliably.

For geo holdouts, sample size is measured in regions. The variance between regions is typically much higher than between individual users, which means you often need 10 or more regions per group to have adequate power. With fewer than five regions per group, results are rarely conclusive.

A quick power check before you launch can save weeks of wasted testing. If the numbers show your test is underpowered, extend the duration, increase the holdout size, or accept a higher MDE rather than running a test you cannot trust.

How to do a quick power calculation

You do not need specialist software for a first estimate. Input your baseline conversion rate, your expected lift percentage, your desired significance level, and your power target into a standard online power calculator. These are widely available for free and require no statistical background to use.

For geo experiments, Alphabet's open-source GeoX tool provides power analysis specifically designed for geographic tests. Meta and other platforms also offer pre-test power analysis in their testing interfaces.

What to do when you cannot get enough sample

If your business is too small for a standard test, you have options. Extend the test period to accumulate more conversions. Accept a higher MDE and design the test to detect only large effects. Or use a marketing mix model as a complement, accepting its inherent limitations without experimental validation.

Running an underpowered test and treating the results as definitive is the worst outcome. It gives you false confidence in a conclusion the data cannot actually support.

Does a larger holdout group always mean a better test?

Not necessarily. A larger holdout reduces the risk of revenue loss from excluded users but does not automatically improve statistical power if the problem is low total conversion volume. What matters is having enough conversions in both the test and holdout groups to detect a meaningful difference.

What MDE should I aim for?

Start with the smallest lift that would change a decision. If you would keep running the campaign at 5 percent lift but cut it at 3 percent, then 5 percent is your MDE. There is no universal standard: it depends on your margins, budget size, and risk tolerance.

How do I know if my test was underpowered after the fact?

If your confidence intervals are very wide, or if the result is not statistically significant despite a sizeable observed difference, underpowering is a likely cause. A post-test power analysis using your actual observed variance will confirm whether the test was adequately sized.

Ready to measure what your marketing actually delivers?

Talk to us