← All posts
Media Mix Modeling

What Is Multicollinearity in Media Mix Modeling?

Vinay Karode · August 6, 2026
What Is Multicollinearity in Media Mix Modeling?

An MMM can predict total revenue accurately and still be too weakly identified to tell you whether Meta or paid search drove it. Reproducing the outcome and separating the contribution of correlated channels are different jobs. When channels rise and fall together, several channel-level stories fit the same revenue history. VIF flags the risk, but no single cutoff tells you a channel ROI is safe to act on. What matters is whether the channel decision survives being rebuilt a slightly different way.

Table of contents

What is multicollinearity?

Multicollinearity is when one channel in the model can be closely reproduced from a combination of the others. It usually happens because media budgets move together: Meta, search, YouTube, and TV all rise for a launch, fall after the holidays, and rise again for the next promotion. The model sees the bundle. Your budget decision asks it to price the pieces.

Take the cleanest version first. Say TV and radio run in identical amounts, in the same weeks, for two years. Revenue rises when both are on and falls when both are off. The data can tell you what the pair did together. It contains nothing that separates the TV effect from the radio effect. Any split you report has to come from an assumption or from evidence outside that history.

That is perfect collinearity, and it is rare. The common case is milder. Channels are not identical, but most of their real movement is shared. The separate effects can still be estimated, but they are weakly identified: small changes to the data or the assumptions swing the split around.

That distinction matters. Near-collinearity does not prove an estimate is wrong. It tells you the data may not hold enough independent movement to pin that estimate down with confidence.

Why can the total look right but the split wrong?

Because reproducing the total and dividing the credit are different tasks, and collinearity only breaks the second. Picture two versions of the same model. One hands most of the incremental revenue to Meta. The other hands most of it to search. When those two channels moved together all year, both versions can fit the revenue history about equally well.

That is why fit and identification have to be judged separately. A low forecast error shows the model reproduces the total under conditions it has seen. It does not show that the split underneath is the only split the data supports. We covered the fit side of this in our piece on out-of-sample R-squared and what a suspiciously perfect fit is really telling you. Collinearity sits one layer deeper, in identification: whether the movement in your data actually supports the specific channel contrast you want to act on.

The problem turns operational the moment you use the model to change the mix. Your history shows what happened when Meta and search moved together. Budget optimization asks the opposite: what happens if Meta goes down while search goes up? The further that sits from the combinations you actually ran, the more the model is extrapolating along a direction it never learned. Google's own Meridian documentation makes the general point, that several models can fit revenue about equally and still recommend different budgets.

You see the instability in the symptoms. Channel ROIs swing between model runs. A channel you know works shows up negative. Rankings flip when you drop a single week of data. When small changes to the inputs produce large changes to the split, the model is not measuring those channels, it is guessing at them.

Why the split is guesswork when the total is not
Drag the credit between two channels that moved together. The combined lift holds. The split, and the budget call, do not.
Meta iROAS
2.40x
swings with the split
Search iROAS
swings with the split
Combined lift
index 100
identified, does not move
all credit to Search even split all credit to Meta
same revenue history, and the budget call is even
Constructed illustration of the extreme case: two perfectly collinear channels with equal, fixed spend. Only the combined lift is identified. The split, and each channel's iROAS, is not, so the same revenue history can justify opposite budget calls. This holds along the historical mix; a reallocation that departs from it is a separate question. Not client data.

What does VIF actually measure?

VIF measures how much a coefficient's variance is inflated because its channel overlaps with the others in the model. Take one channel, predict it from all the rest, and read the R-squared of that side regression. Then VIF = 1 ÷ (1 − R-squared). If the other channels explain 80% of a channel's movement, its VIF is 5. If they explain 90%, its VIF is 10.

In a plain linear regression, a VIF of 5 means that coefficient's variance is five times what it would be if the channel stood on its own. The error bar you actually see, the standard error, widens with the square root of that, so it is a bit over twice as wide. That is exact for the classical case, and it is the intuition worth carrying.

But a modern MMM is not a plain linear regression. It applies adstock and saturation transforms, estimates effects across geographies, and regularizes the result with priors. VIF stays useful as a redundancy check on a specific design matrix. The classical number is not automatically a literal multiplier on the ROI interval in that larger, nonlinear model.

That leads to a question most buyers never ask. Which design matrix was the VIF computed on? Raw spend, or the adstocked and saturated media the model actually fits? A national table, or geo-week observations? Meridian computes VIF on the transformed, scaled variables the model uses, not on raw spend, so a table built on raw spend describes numbers the model never sees.

How VIF inflates a coefficient's error bar (classical case)
Drag the slider. The estimate holds still. The error bar widens with the square root of VIF.
Others explain
0%
of this predictor's movement (R squared)
VIF
1.0
below screening lines
Error bar vs independent
1.0×
wider (scales with √VIF)
0% (independent) 80% (VIF 5) 90% (VIF 10) 95% (VIF 20)
VIF = 1 / (1 − R squared), and in a textbook linear regression a coefficient's standard error widens with the square root of VIF. Both are exact for that classical case. A real MMM is nonlinear and often Bayesian, so treat this as intuition for what collinearity does to precision, not a literal multiplier on an iROAS posterior. The 5 and 10 marks are conventional screening lines, not pass or fail. The estimate and baseline interval are illustrative, not client data.

What is a good VIF score?

There is no universal score that makes an MMM valid or invalid. Values of 5 and 10 get used as screening lines, and they are fine as prompts to investigate. They are not natural laws. O'Brien's widely cited analysis of VIF rules of thumb warns specifically against treating 10 as an automatic failure. What a given VIF costs you depends on the sample size and on how much independent movement the channel actually has.

The scale marketers assume here is often wrong by orders of magnitude. Google's Meridian sets its extreme blocking threshold, the point where it stops the model, at a VIF of 1000, not 5 or 10, and lets you tune it. That number is meant to catch near-perfect redundancy and numerical breakdown, not to certify that everything below it is ready to move budget on. Meridian is blunt about the mechanism too: high multicollinearity widens the credible intervals of the coefficients and makes the result less reliable.

A higher VIF means a channel carries less independent information than the others. Whether that is a problem depends on how uncertain and how stable the estimate is, and on the decision in front of you. A wobbly coefficient may be fine for a directional plan. It is not fine when the call is to move several million dollars between two channels whose separate effects are mostly assumption.

What can't VIF tell you?

VIF covers one source of instability, and only one. A low VIF does not make a channel ROI causal. The model can still miss a confounder, control for something it should have left alone, use the wrong curve shape, or misjudge how long an effect lingers. Whether it controlled for everything that drives both spend and revenue is an assumption you cannot prove from the data, so a model can carry a low VIF, a strong fit, and a biased channel estimate all at once.

A clean VIF cannot rescue that. Recent research from Wharton and London Business School, "Your MMM is Broken," shows that even nonlinear and time-varying effects are often not separately identifiable from standard marketing data, especially when media is autocorrelated, and that the fix is designed experiments, not a better fit statistic.

VIF is a clue. Identification is the case you have to build.

What should you ask your vendor?

Do not stop at "what is the VIF for each channel." Ask the harder version: show me which channel decisions stay the same under reasonable alternative models, priors, and data. A defensible answer covers four things.

  1. The exact diagnostic context. Which variables, transformations, geographic level, and time level was the VIF computed on. A single national VIF table on raw spend is not a diagnosis for a geo-level model built on transformed media.
  2. The trade-off between the correlated channels. Ask to see how the two channels' estimates move together in the posterior. When one channel's effect rises every time the other's falls, the model is showing you a substitution it cannot resolve. The marginal numbers alone hide it.
  3. Prior sensitivity. For a Bayesian MMM, ask to see the prior and the posterior together, then a rerun under other defensible priors. A stable posterior is not proof the data identified the effect. A strong prior can steady an answer by supplying what the data did not. The vendor should be able to say what came from the data, what came from the prior, and whether the decision changes under another reasonable prior.
  4. Specification and decision stability. The real test is not whether the coefficient moved. It is whether the action moved. Ask what happens to the channel ROI and the budget call when the vendor makes reasonable changes to adstock, saturation, controls, lag, or the estimation window, and when the reallocation is run across the plausible model set. If one specification says move budget from Meta to search and another says the reverse, the model has not earned that call, and the honest move is to report the disagreement rather than average it away.

This is check four of the nine questions we give clients for grading a measurement vendor. It is one of the harder ones to dodge, because the answer is either a specific number and a specific decision or it is a change of subject.

How do you fix multicollinearity in an MMM?

There is no single fix, because different problems hide under the same diagnostic. Work through them in order.

  1. Correct the specification first. Some collinearity is self-inflicted. Two variables may stand in for the same activity, a channel may appear as both spend and impressions with no reason, or a trend term may swallow most of a slow-moving channel. Decide whether each variable earns its place on causal and business grounds before you touch anything. Dropping or combining variables just to lower VIF can introduce bias or erase a distinction the business needs. O'Brien makes this point directly: the automatic cures for high VIF often cost more than the collinearity did.
  2. Create independent movement. The most durable source of channel identification is data where the channels vary on their own. Casually nudging search up one month and social the next does not count, because that sequence can still be confounded by seasonality, promotions, or expected demand. The stronger route is a designed intervention: randomized geo budget changes or staggered launches that create clean contrasts on purpose. The "Your MMM is Broken" research reaches the same conclusion, that strategically manipulating spend is what pins down effects the historical record cannot separate.
  3. Calibrate with experiments. An incrementality test supplies information the observational history is missing. In a Bayesian MMM, that experimental read becomes a channel-specific prior, so a tangled channel is disciplined by evidence instead of guesswork. This is the core of a calibrated MMM, and why we do not treat modeling and experimentation as separate products. Google's Meridian calls incrementality experiments perhaps the strongest basis for a prior, while warning that the translation is not one-to-one: the experiment's uncertainty, and any gap in geography or spend level, has to carry through. Calibration does not make the historical variables less correlated. It adds outside information that helps separate them, which should be disclosed, since part of the answer now comes from the experiment rather than the MMM data.
  4. Use regularization, and show what it is doing. Ridge, shrinkage priors, and sign constraints can steady unstable estimates and rule out impossible values. That is useful. It is not the same as recovering information the data never held. With weak priors, correlated effects can stay broad and hard to sample. With strong priors, the result can turn stable but prior-driven. The question is not whether the model is Bayesian. It is which assumptions are steadying the result, and whether the recommendation survives another reasonable set of them.
  5. Report channels jointly when that is what the data supports. If Meta and YouTube moved together, their combined contribution may be solid even when the split between them is guesswork. Reporting them as a pair, with a wider shared interval, is more honest than two precise-looking numbers. But "digital drove this much" is no help when the next decision is Meta versus YouTube, so joint reporting marks the limit of what today's data supports. It does not remove the need for a test that separates the two.

What standard should your vendor meet?

A high VIF is not a failing grade, and a low one is not a passing grade. The failure is a channel number carried with more confidence than the data and the assumptions earned.

So the bar is not whether the model produced a number. It is whether it earned the contrast behind the decision. If you want a partner who shows you where the model is identified, where it is leaning on assumptions, and what evidence would close the gap, book a scoping call and we will walk your own model's numbers with you.

FAQ

What VIF value is too high for an MMM?

There is no universal cutoff. Above 5 or 10 are conventional screening signals, not automatic failures, and the exact number depends on the design matrix, the channel's independent movement, sample size, model structure, and the decision at hand. For scale, Google's Meridian sets its extreme blocking threshold at a VIF of 1000. The practical question is whether the channel estimate and the recommendation stay stable.

Can an MMM have a strong R-squared and still be unreliable at the channel level?

Yes. Fit measures how well the model reproduces total revenue. It says nothing about whether the data uniquely supports the channel split underneath. Several models can fit revenue similarly while assigning materially different effects to correlated channels, which is exactly the number you move budget on.

Does a Bayesian MMM solve multicollinearity?

No. Priors can steady correlated effects and express uncertainty, but the framework cannot manufacture independent movement the data never had. Informative priors or experimental calibration can add outside information, though the result may then lean materially on that evidence, which is why prior-sensitivity analysis matters.

Why do my channel ROIs change between model runs?

Multicollinearity is one cause, not the only one. Instability can also come from limited variation, a few influential weeks, prior sensitivity, misspecification, confounding, or weak identification of the adstock and saturation parameters. Treat it as a symptom to diagnose, not proof of a single cause.

What is the difference between correlation and multicollinearity?

Correlation is between two variables. Multicollinearity is a many-variable property: one channel can be reproduced from a combination of several others. A model can have serious multicollinearity even when no single pairwise correlation looks extreme, which is why VIF exists to catch the broader redundancy.

Should highly collinear channels be combined?

Only when the combined effect is more stable, the grouping means something to the business, and the decision does not require separating the channels. Combining variables just to lower VIF hides the identification problem rather than solving it.