MMM calibration uses the results of controlled experiments to constrain or adjust the model's channel coefficients. When your model's estimates align with your experimental results, you have much stronger confidence in its budget recommendations.
Marketing mix models are powerful but imperfect. They estimate channel contributions from observational data, which means they can confuse correlation with causation. A channel that runs heavily during periods of high organic demand may appear more effective than it really is. Incrementality tests correct for this.
What calibration means in practice
Calibration means using the output of an incrementality test, specifically the estimated lift and iROAS, as a prior or a constraint when fitting the MMM. You are telling the model: "we know from an experiment that this channel produced approximately X percent lift. Make sure your estimates are consistent with that."
This is especially valuable for channels where observational data is noisy or where the channel correlates strongly with organic demand, making the model prone to over-attribution.
The calibration workflow
- Run an incrementality test for one or more key channels and obtain a reliable lift estimate with confidence intervals
- Express the experimental result as an iROAS range compatible with the model's output units
- Set a Bayesian prior or a constraint in the MMM that penalises coefficient values inconsistent with the experimental range
- Refit the model with the calibration applied and check whether the channel coefficient moves into alignment
- Repeat for additional channels over time to improve overall model accuracy
Bayesian MMMs and calibration
Bayesian marketing mix models, such as Meridian (Google's open-source framework) or Robyn with Bayesian extensions, are particularly well-suited to calibration because they allow you to formally set priors on channel effectiveness based on experimental evidence.
In a Bayesian framework, the experimental result becomes a prior distribution on the channel coefficient. The model uses this prior along with the observational data to produce a posterior estimate that is constrained by what the experiment found.
You do not need a Bayesian MMM to calibrate. Frequentist models can also be constrained using experimental results as bounds on channel coefficients. The constraint is less mathematically elegant but produces similar practical improvements in model accuracy.
What happens when calibration reveals a large discrepancy
If the model's uncalibrated channel estimate is far from the experimental result, that discrepancy is valuable information. It usually means one of three things: the model has a specification problem, the experimental test was flawed, or there is a genuine interaction effect the model is not capturing.
Investigate the discrepancy before simply forcing the model into alignment. If the experimental design was sound, trust the experiment over the model. If there are concerns about the test validity, re-run the experiment before making large budget decisions.
Building a calibration programme
One calibration test is better than none, but a systematic programme of tests across channels is the gold standard. Prioritise channels by spend and by how uncertain the model's estimates are for those channels. High-spend channels with wide model uncertainty should be tested first.
Aim to run at least one calibration test per major channel per year, and more frequently for channels where spend or strategy changes significantly.
Do I need to refit the MMM after every incrementality test?
Not necessarily after every test, but you should update the model whenever a calibration test produces results that are meaningfully different from the current model estimates. Treat the MMM as a living model that improves as more experimental evidence accumulates.
Can I calibrate an MMM with platform-run holdout tests?
Platform-run holdout tests can be used as calibration inputs, but they tend to produce higher lift estimates than independently run geo tests. If you use platform results, apply a conservative adjustment or weight them less heavily than geo-based experimental results.
What if I only have one incrementality test result to calibrate against?
A single test is a meaningful improvement over no calibration. Use it as a constraint on the channel it tested. For other channels, you must rely on the model's observational estimates until you have experimental data. Be transparent about which channels are experimentally validated and which are not.
