Do YouTube ads actually work? What 92 incrementality tests show

Often, yes, but your in-platform reporting is the wrong place to check. Across 92 YouTube and Demand Gen holdout tests, the median incremental return was 2.01x, about two dollars of incremental revenue for every dollar of test spend. That is a causal number, not an attributed one, and the two routinely disagree. So the real question is not which platform got credit. It is whether the sale would have happened anyway.
Your in-platform YouTube reporting can be completely accurate and still send you the wrong way. Google can measure clicks, engaged views, and view-through conversions, so the channel is not invisible to modern attribution. But assigning a conversion to an observed ad interaction is a different thing from estimating whether that conversion would have happened without the ad. One is bookkeeping. The other is cause.
That gap matters, because budget decisions are causal decisions. You do not really care which platform got credit for the sale. You care whether the sale would have happened if you had never spent the money.
The 2.01x figure below is not "YouTube's ROAS." It is the median across a specific set of tested interventions, with the middle half of studies landing between 1.73x and 2.50x, across $2.3 million in test spend. Small difference in wording. Large difference in what you can honestly claim.
What does each measurement method answer?
Four methods, four different questions. Attribution observes which tracked interactions happened around a sale. Surveys capture what customers say influenced them. Incrementality estimates what the advertising actually caused. Allocation asks what the next dollar will do. They disagree because they are built to answer different things, and the disagreement is usually the useful part.
Most measurement arguments are really question problems. People line up attribution against surveys against incrementality against MMM and fight over which one is right. That is the wrong fight.
Attribution observes behavior. It tells you which measurable interactions happened around a conversion and assigns credit under whatever rules you set.
Post-purchase surveys capture reported influence. They show what customers say introduced them to you, including touches the tracked path misses.
Incrementality estimates cause. It asks what would have happened to the business if the advertising had not run.
Allocation needs a response curve. It asks what happens to the next dollar, which is a different question again.
A channel can have weak attributed ROAS and strong incremental lift. A survey can catch an influence no tracked path ever recorded. An experiment can show a channel is incremental at the tested spend without telling you whether doubling the budget is smart. The mistake is asking one method to answer all four questions.
Why is YouTube such a clear example?
Because YouTube's influence usually lands before the sale, not on it. Someone sees your ad Monday. They do not buy. On Thursday they search your brand and purchase. You can observe pieces of that, an engaged view here, a search click there. What you cannot observe is what that person would have done on Thursday if Monday's ad had never run.
That missing outcome is the counterfactual. You cannot watch the same customer in both worlds, the one where they saw the ad and the one where they did not. Attribution assigns credit among the events you observed. Causal measurement tries to estimate the outcome you did not. That is the whole problem, and no amount of extra touchpoints in an attribution report turns one into the other.
What does the survey data show?
That discovery and tracking disagree constantly. In the Q2 2026 YouTube Ads Report, Fairing looked at roughly 1.7 million orders where buyers named YouTube as where they found the brand. Only 1.97% carried a YouTube UTM. About 65% carried no UTM at all, and many of the rest were tagged to channels like Search, Meta, or Email.
Be careful with what that proves. It does not prove YouTube caused all 1.7 million purchases. A survey is not a causal estimator. People misremember, simplify, and read "how did you hear about us" in their own way, and a missing YouTube UTM does not mean there was zero YouTube signal anywhere in the stack.
What it shows is narrower and still striking. What customers report as their discovery source and what the observable UTM path records pull apart, at scale. That tells you where to investigate. It does not tell you what the channel is worth. For that you need a different design.
What does a YouTube holdout really measure?
A holdout changes the advertising on purpose and watches whether the business changes with it. YouTube keeps running in one set of markets and is held out of comparable ones, or a live channel is paused in selected markets and left on elsewhere. Then you compare outcomes against the markets you did not touch.
This is incrementality testing, and it is not magic. You still never observe the same market treated and untreated at the same moment, so the untreated outcome is a counterfactual you have to estimate. Whether that estimate holds up depends on the test. Were the markets comparable. Was there enough spend to detect an effect. Did something else hit one group during the window. Did people cross between test and control. Did the pre-period model actually fit. Was the test long enough. Those are not footnotes. They are the difference between a number you can take to your board and one that wasted a month.
Experiments are powerful because you deliberately create variation in exposure instead of trying to recover cause from correlations that happened on their own. That does not make every experiment right. It makes the causal question answerable in a way plain observation usually cannot. There is a clean external example. Gordon, Zettelmeyer, Bhargava, and Chapsky compared observational estimates against 15 large Facebook experiments covering roughly 500 million user-experiment observations, and even with that much user-level data, the observational methods overstated the true effect by roughly three times in about half of them (Gordon et al., Marketing Science, 2019). More data did not fix it, because the limit was never data. It was the question the design could answer.
What did 92 YouTube holdouts find?
A median study-level iROAS of 2.01x, with the middle half between 1.73x and 2.50x, across $2.3 million in test spend. The set breaks down into 69 geo holdouts and 23 inverse holdouts, and in 86% of studies the 90% confidence interval excluded zero. Tests used 120-day pre-test windows and multiple model families, with validation and QA on every counterfactual.
And yes, some tests came back below 1x. Good. That is what a real dataset looks like. A channel losing money is not a failed experiment. If the test was valid, discovering that a treatment destroyed value is the experiment doing its job. I would be far more suspicious if 92 tests somehow produced 92 winners.
The fit diagnostics, a median R-squared of 0.817 and a median MAPE of 0.124, describe how well the control model reproduced the untreated business before anyone looked at the treatment effect. Good pre-period fit is necessary, not sufficient. It does not prove the effect estimate is unbiased, which is why you pick the model before you look at the treatment effect.
What does a 2.01x iROAS mean?
Less than the headline suggests, and that is the point. These 92 studies are not a random sample of every advertiser running YouTube. They are brands that chose to test YouTube with Stella, so the group has selection effects, some advertisers appear more than once, and the thing being summarized is the study, not YouTube as a whole.
So the honest conclusion is not "YouTube produces a 2.01x incremental ROAS." It is "across these 92 tested interventions, the median measured incremental ROAS was 2.01x." That distinction limits how far the number travels. It says what happened in this sample. It does not promise what YouTube returns for another advertiser, creative, audience, spend level, or moment in time.
The right response to this dataset is not "YouTube works." It is "YouTube deserves to be tested causally before it gets cut on attributed ROAS alone."
How solid is the 2.01x result?
Solid in the way that matters: it barely moves when you change how you summarize it. The median across all 92 studies is 2.01x. Pool incremental revenue against test spend and you get 2.09x. The geo subset lands at 1.99x, the inverse subset at 2.04x. Same story, four different cuts.
Be careful not to oversell it. These are not four independent replications. The median and the pooled figure summarize overlapping studies, and the geo and inverse numbers partition the same dataset by design. Their agreement does not prove the effect four separate times. What it does show is that the number does not depend on one weighting choice or one holdout direction. Change the summary, split the designs, and the headline barely moves.
Now a word on statistical significance, because it gets misused constantly. A test reads 99% significant, everyone high-fives, and that is not enough. Significance tells you how unlikely your estimate would be if the true effect were zero, given the model's assumptions. It does not tell you the estimate is exactly right, that the interval is narrow, that the effect is economically meaningful, or that the design was sound. A highly significant result with an interval too wide to budget against is still useless, and a result that misses a threshold is not proof the channel did nothing. Look at the interval, the design, the counterfactual, and the size of the decision.
Can one holdout guide budget scaling?
No. Say a clean holdout shows $100,000 of YouTube spend generated $200,000 of incremental revenue. You have learned something real about that intervention. You have not learned that the next $100,000 returns 2x. A single experiment estimates the effect of the treatment you tested, not the marginal return at a different spend level.
Returns diminish as spend climbs. Audiences saturate, frequency rises, creative fatigues, auctions shift, seasonality moves. Allocation asks for the marginal return at other spend levels. That is a different question.
To get at the curve you run experiments repeatedly, vary treatment intensity where you can, and use those causal anchors to discipline a response model or media mix model rather than letting a model infer cause from history. An experiment gives you an anchor. Allocation requires a curve. That is why incrementality should run as an operating system, not an annual validation exercise.
Should your measurement methods all agree?
No, and be suspicious when they do. Attribution looks at observable behavior. Surveys look at remembered influence. Experiments look at what a treatment changed. MMM estimates how outcomes move across channels, spend, and time. They examine different evidence, so they will not land on the same number. The disagreement is the information.
If customers keep naming YouTube and attribution barely sees it, that is information. If attribution loves branded search but a holdout shows most of those sales happen anyway, that is information. If an experiment says YouTube is incremental at $100K but your response model says returns collapse at $300K, that is information.
The goal is not to force every method to agree. It is to understand why they disagree and which evidence answers the decision in front of you. Attribution cannot become causal because it holds more touchpoints. A survey cannot become causal because the sample grows. A holdout cannot become a budget optimizer because it produced a precise effect. An MMM is not ground truth because its R-squared is high.
What should this change about your measurement?
Stop treating attributed performance as the verdict. The finding here is not "spend more on YouTube." It is that attributed and causal performance can diverge enough to flip a budget decision, and it cuts both ways. You can keep funding a channel that looks great in attribution and barely moves the business, or cut one that looks weak in attribution and is quietly doing more than the platform credits. YouTube is a clean example of the second.
Your attribution is not ground truth. Neither is your survey, your holdout, or your MMM. They are different forms of evidence answering different questions, and the discipline is knowing which question you are asking. What happened: use observational data. What might that journey be missing: ask the customer. What changed because the advertising changed: run a causal test. Where the next dollar should go: estimate the response curve, anchored by experiments wherever you can.
The observable isn't the incremental. And the incremental isn't necessarily the marginal. In plainer terms: observation is not causation, and causation is not allocation. Once you separate those three, most of the measurement arguments marketers have had for a decade get a lot easier.
The 92 studies here are Stella's contribution to the Q2 2026 YouTube Ads Report, alongside Northbeam, Fairing, and Taikun Digital. If your in-platform reporting says YouTube does not work, do not assume the reporting is broken. Ask the better question: what would have happened if you had not spent the money. That is the number worth finding.
Frequently asked questions
Does this study prove YouTube works?
No. It shows that across 92 YouTube and Demand Gen interventions tested with Stella, the median study-level incremental ROAS was 2.01x. The advertisers were not randomly sampled, so the number should not be treated as a universal YouTube benchmark. It is evidence that YouTube can produce meaningful incremental revenue and that attributed performance alone may be insufficient to judge it.
Is incrementality testing always more accurate than attribution?
No. They answer different questions. Attribution describes observed journeys and assigns credit. Incrementality estimates what changed because advertising changed. A badly designed experiment can still produce a bad causal estimate; its validity depends on treatment assignment, power, counterfactual quality, spillover, and other design assumptions.
What if my holdout is not statistically significant?
A non-significant result does not automatically mean the channel had no effect. A holdout can be inconclusive for several reasons: the true effect may be small or zero, the test may be underpowered, the outcome may be noisy, or the design may have contamination or weak controls. Start with the confidence interval and the diagnostics. A narrow interval around zero tells you something very different from a wide interval spanning large positive and negative effects.
Does a 2x incremental ROAS mean I should double spend?
No. An experiment estimates the effect of the treatment actually tested. It does not identify the marginal return on the next dollar. Scaling decisions require evidence about the response curve, from repeated experiments, changes in treatment intensity, and models calibrated against causal measurements.
Why use post-purchase surveys if they are not causal?
Because measurement is not only causal estimation. Surveys surface influences and discovery sources that behavioral tracking underrepresents. They tell you where attribution may have a blind spot, and therefore where a causal test is worth running.
Does the same logic apply to CTV and linear TV?
Yes at the conceptual level: observing an exposure and assigning it credit are different from estimating what it caused. But CTV and linear experiments have their own power, spillover, geography, and design requirements. The 2.01x result here applies to the YouTube and Demand Gen interventions in this dataset, not to CTV or linear.