← All posts
Incrementality

What does a good incrementality test design look like?

Vinay Karode · September 28, 2026
What does a good incrementality test design look like?

A good incrementality test design is one that could have produced a real number before a single ad turned off. For a geo holdout, Stella runs four diagnostics in advance: the test was feasible at your spend, the test markets moved with the control markets historically, a placebo test on the pre-period found no material lift, and a held-out slice of history could be predicted. None of them proves the result is causally valid. They tell you whether the counterfactual behaved credibly before treatment, and that is the part you can check. If a geo-test vendor cannot show you how it checked each of those, the iROAS is an opinion with a decimal point.

Table of contents

Two vendors can run the same channel in the same markets for the same six weeks and hand you different iROAS numbers. One checked the design before launch. The other randomized a handful of DMAs, skipped the balance check, and hoped. Both will have a confidence interval. One of them was designed to give the counterfactual a fighting chance.

That is why test design is Check 6 of the nine questions a measurement vendor should be able to answer: what was the design, how was the control built, what was excluded, and what could have biased the result? A number is not auditable. The way it was made is. Passing the checks below makes the work auditable. It does not make the conclusion right, and any vendor who tells you otherwise is overselling.

Was the incrementality test feasible at your spend?

Before you pick a single market, answer one question: can this test see a lift big enough to change your decision? Every test design has a smallest lift it can reliably detect. If that is bigger than the lift you would need to act on, the test cannot give you an answer, however well the markets are chosen. That is why feasibility comes first. It decides whether the test runs at all.

The question we hear most is "what's the minimum spend?" Our rule of thumb is $80K per channel per month. It is a starting point, not a guarantee, because it is based on sales volume. Whether a test can see anything depends on how much volume the markets carry and how much it moves week to week, how many geographies you have, how long the test runs, and how big the spend change is. Stella runs that calculation before every test, and sometimes the answer is no.

No is the right outcome more often than people like. A test whose smallest detectable lift is bigger than the effect you care about costs you the holdout period, the fee, and a "no lift" readout that gets read as "the channel does not work." Our post on minimum detectable effect covers why those are not the same conclusion.

Ask your vendor for the feasibility number. If they did not calculate one, they did not know whether the test could work before they ran it.

Did the test markets move with the control?

Correlation is the first of three location checks that run before anything turns off. Test and control markets need to have risen and fallen together over the pre-period, week by week, for months. If they did not track before the test, any gap that opens during it could be your ads or could be the gap that was already forming.

Most tests do a version of this. They eyeball a chart, pick markets that feel similar on size and region, and call it matched. Feeling similar does not count. Stella's platform scores every proposed pairing and flags anything under 0.60. That is an operational screening threshold from our own experience, not a statistical law. A weighted donor pool can build a good counterfactual out of markets that only correlate moderately on their own, and two markets can correlate at 0.95 just by sharing a season and still make a poor pair.

Markets do not need identical revenue levels. What matters more is whether the control can reproduce the test group's week-to-week shape. A test region doing $400K a month and a control doing $90K can be a fine pair if their weekly shapes match. If test markets were climbing while controls were flat, the test reports a lift that is really the climb continuing. And the window has to be long enough to mean something. Stella typically builds synthetic controls from around 120 days of pre-period sales, and adjusts the window when seasonality, promotions, or a structural change make a different period a better fit.

Correlation gets you to a short list. The next two checks tell you whether the proposed counterfactual survives harder scrutiny. Ask for the number anyway. A vendor who ran the check has one. A vendor who says "the markets were comparable" has an adjective.

Does a fake test find a fake lift?

The placebo test is the second location check, and the one that catches the most. You may hear it lumped in with A/A tests, a related but different procedure. Take a stretch of the pre-period when nothing changed. Run the full analysis on it as if a test had happened. If the model reports a material lift where nothing was assigned, the counterfactual is not stable enough to trust until you know why. It could be finding effects in noise, missing a shock that hit one group, or leaning on control markets that drift. A credible design should not keep producing placebo effects big enough to matter for the decision the real test is meant to inform, and it has to hold across more than one window, because a single window can pass by luck.

It is the closest thing a geo test has to a smoke alarm. It catches what correlation misses: markets that track on average but drift apart in ways the synthetic control cannot follow. A promo that hit one region a week early. A store opening.

Meta's open-source GeoLift package bakes a version of this into market selection. For every candidate set of test markets it simulates a 0% lift and reports what the model found anyway, and its walkthrough says good selections have that number very close to zero. Nothing exotic. People just skip it, and this is the one that bugs me, because it takes an afternoon and it is the check that would have caught a lot of the bad readouts I have seen.

A good answer: "We ran placebos on the eight weeks before the window. Estimated lift stayed well under the smallest lift the test was built to detect." A weak one: "The model was validated." Validated on what?

Can the past predict itself?

The third check is the 80/20 predict. Hide the last 20% of the pre-period, build the synthetic control from the first 80%, and see how well it predicts the slice it never saw. A model that predicts history it did not study is less likely to be fitting noise than one graded only on its training data. Same logic as out-of-sample validation for an MMM.

The held-out 20% stays untouched during market selection and model fitting. If you use it to tune the model and then report its error, it is no longer out of sample. The number to ask for is the prediction error on that slice. Stella reports it as MAPE, the average percentage miss, as one readable diagnostic alongside the others. It is a first screen, not the whole read, since MAPE misbehaves near zero. In our 225-test benchmark, models under 15% MAPE on the holdout were more often able to distinguish an estimated lift from zero, so we use 15% as an internal selection heuristic. It is not a universal validity line. And reaching significance is a statement about precision, not accuracy. A tighter fit shrinks the noise around the estimate, which makes a clear read more likely. It does not prove the estimate is unbiased. That is the placebo test's job.

The lift is the gap between actual sales in the test markets and what the synthetic control predicted, so pre-period prediction error is a practical read on how much unexplained movement the counterfactual carries into the test. If the estimated lift is small next to the counterfactual's unexplained error, the interval can easily come back too wide to act on. You get a number and no decision.

Only after the design clears all three diagnostics, on top of feasibility, does the channel turn off. That earns the test the right to run. It does not settle what happens during the window.

Stella's benchmark of 225 geo-based incrementality tests is a self-selected set of DTC brands that chose to test, so treat it as a sample and nothing wider. Within it, pre-period fit predicted whether a test reached significance more strongly than budget or duration did. Part of that could be stable businesses being both easier to fit and easier to detect lift in. But among the things we measured, pre-period fit, the part design influences most, had the strongest observed relationship with significance. Full benchmark here.

How do I audit my vendor's incrementality test design?

Ask the Check 6 question in full: what was the design, how was the control built, what was excluded, and what could have biased the result? Start with the design, and ask for four things by name: the feasibility calculation, the pre-period correlation, the placebo results, and the prediction error on the held-out window. A vendor who ran the checks has all of it on hand. A vendor who says "trust the result" did not run them, or ran them and did not like what they found. I never trust the first number I get, including from our own platform. The checks are how I earn the right to trust the second one.

CheckWhat it catchesWhat failure looks like
FeasibilityA test too small to see the effect"No lift" that really means "couldn't see it"
CorrelationTest and control markets that never trackedA lift that is really a pre-existing trend
Placebo testA model that finds effects in noiseA confident number on a window where nothing happened
80/20 predictA control that memorized instead of learnedA lift number sitting on a wide prediction error

Next, how the control was built. A real incrementality read makes two comparisons: test markets before versus after, and test markets versus a control built to mimic what they would have done on their own. If the tool reports test minus control during the window and stops, on markets that were hand-picked or never balanced, it credits your channel with every difference between the two groups, including the ones that existed before the test. With many well-randomized, balanced markets a simple difference can be a fine estimator. Most geo tests do not have that luxury. Difference-in-differences nets out the pre-period gap and works when the two groups plausibly share a trend. Synthetic control is for when no simple control group reproduces the test markets well and a weighted blend of them can. Neither is universally better. Unadjusted subtraction on unbalanced markets is the one to push back on.

So ask which method they used and why it fit these markets. A vendor who picked synthetic control because a plain average did not reproduce the pre-period is doing the work. One who picked it because it is the default is not.

Then what was excluded, which is where geo tests live or die. Commuters cross DMA lines, national campaigns and TV markets reach both groups, and word of mouth does not respect a holdout. Ask which markets were left out and why, and whether anyone verified the contrast actually happened: delivery should move in whichever group the design changed, in impressions and not just spend, and hold steady in the other. The more exposure bleeds between the two groups, the smaller the real contrast and the harder the result is to read.

Last, what could have biased the result. Leakage is one answer. An unmodeled promo concentrated in the test markets is another, since it is indistinguishable from your ads. So is running test and control in different weeks, which counts seasonality as lift. A vendor who did the work can tell you what might have moved the number and what they did about it.

None of this requires a statistician in the room. It requires asking for the numbers by name and noticing which questions get a straight answer and which get a subject change. Our guide to vetting a measurement consultancy applies the same approach to the whole engagement, and once you have a number you trust, the next question is what to do with it.

This is how Stella grades every geo test we run: feasibility first, three location checks before anything turns off, two comparisons at the end. Every number comes with the design attached, so you can check our homework. To see it on your own channels, book a demo.

Frequently asked questions

How many markets does a geo incrementality test need?

No fixed number. Power comes from how many geographies you have and how much they vary week to week, plus the size of the spend change and the test length. A brand with 15 well-correlated DMAs and a big spend change can be better powered than one with 40 noisy ones.

Can an incrementality test pass all four design checks and still be wrong?

Yes. The checks make the design auditable. A well-designed test can still be thrown by something that happened only during the window. A confidence interval captures statistical uncertainty. It does not protect you from an unmeasured confound.

What is a synthetic control in plain English?

A weighted blend of your control markets, built so its combined history matches the test markets' history as closely as possible. During the test the blend keeps running with no ad change. The gap between what the test markets did and what the blend did is the estimated lift.

Why not just randomize which markets go in the test?

Randomization is not the problem. Randomizing a handful of very different geographies without blocking, a balance check, or a power analysis is. Randomization helps. But with a small number of very different geographies, we still run the pre-period checks to confirm power, balance, and precision before launch.

Does incrementality test design apply to MMM too?

No. Test design is an incrementality check. MMM has its own set, starting with out-of-sample prediction error, covered in the MMM posts in this series. The one idea both share: a model has to be graded on data it never saw.