← All posts
Incrementality

How Often Should You Run Incrementality Tests?

Brenden DelaRua · April 28, 2026
How Often Should You Run Incrementality Tests?

Run at least one incrementality test every month, and keep one live year round. Not every channel at once, but always something in market. A test is a causal read on one channel, at one spend level, in one competitive moment, and you reallocate budget every month. The reads under those decisions should be about as recent as the decisions.

Table of contents

How often should you run incrementality tests?

At least one test completing every month, rotated across channels rather than all at once. A result has no fixed shelf life; it decays as fast as your creative, competition, and seasonality move, often within weeks. Keep the evidence about as fresh as the budget decision it supports.

Two different problems make a single number unreliable. Lift varies enormously between channels, brands, and spend levels, so you cannot borrow another test's result or assume last quarter's still holds. Your own channel also drifts as delivery re-optimizes and rivals shift bids. The data below shows the first; the next section is the second.

Across 225 geo incrementality tests on Stella's self-service platform, a self-selected set of brands that chose to measure rather than a market average, the median iROAS was 2.31x, with the middle half from 1.36x to 3.24x. That spread is between tests, not within one channel over time, which is why a number from one account rarely transfers to another. The full breakdown splits it by channel and spend tier.

Why does one test stop being true?

Its estimate is tied to the conditions inside the test window. Creative fatigues, competitors change their bids, seasonality shifts demand, and the platform re-optimizes delivery toward a different slice of your audience. A result from six months ago carries none of that, so it slowly describes a market that has moved on.

Platform-reported attribution will not flag the drift, and it starts from the wrong baseline. Comparing 15 large Facebook experiments against their observational estimates, Gordon et al. found the observational methods overstated lift, often by around 3x, because ad delivery is not random. Only a fresh experiment resets the baseline.

What is always-on incrementality testing?

A holdout or geo split you keep running continuously, so the estimate refreshes as conditions change instead of expiring between one-off studies. It carries an ongoing revenue cost, and it has to be sized for statistical power and read with sequential methods, or repeated peeking turns noise into false alarms. Run it properly and it flags a fading channel months earlier than a periodic test.

For a brand where one channel misfiring costs six figures a quarter, that earlier warning is worth the drag. Stella runs this as always-on incrementality.

Is a holdout test worth the cost?

Almost always. A holdout withholds ads from a control group, so that group converts less and you forgo some revenue during the test. That cost is real, and it is small next to the spend the result governs. A holdout cannot prove a channel does nothing. It can bound the channel's lift, and a bound that sits below what you need to break even is enough to act on.

Say a channel takes $500,000 a year and a test costs roughly $40,000 in withheld revenue. If the result puts its lift under your breakeven, you stop funding it at that level and move the money to a channel you have measured higher. You spent $40,000 to redirect $500,000, not to recover it. The Performance Max holdout walkthrough covers running one without leaking the control group.

Geo experiments or holdout tests?

Match the method to the question. A geo experiment suits channel-level or brand spend and new channels: you split regions into test and control, usually with matched markets or a synthetic control rather than raw randomization, because a handful of unequal geos randomizes poorly and leaves you underpowered. A holdout suits campaign or audience-level questions where you have solid conversion coverage.

Geo experimentHoldout test
What it splitsRegions, matched or synthetic controlA user-level group inside a channel
Best forChannel or brand spend, new channelsCampaign or audience-level questions, retargeting
NeedsEnough comparable geos, clean separationPlatform-side identity, solid conversion coverage
Main riskContamination and underpowered geo selectionOverlapping tests, a leaky holdout

Setup for each is in the geo-testing guide. A geo test needs no user-level identity; a holdout leans on the platform's own identity graph, not third-party cookies. Both survive signal loss better than pixel attribution, the case for incrementality over last-click.

How do you build a monthly testing loop?

Six steps, repeated. One budget question, a decision rule written before results, the right method, a clean design, the decision, then the next question. What compounds is not any single test but a monthly-improving map of your channel mix.

Step 1: Pick one specific budget question. Not "does paid social work," but "does Meta prospecting generate incremental lift at our current $50K per month?"

Step 2: Write your decision rule before the test starts, so the result cannot be rationalized after you see it. iROAS comes back as an interval, and you read it against two lines, not one: zero, which asks whether the channel did anything, and your breakeven iROAS, which asks whether it paid.

Where the interval sitsWhat it meansMove
Entirely above your breakeven iROASClears the hurdle at the tested spendScale, near that spend level
Straddles breakevenReal revenue, profitability unresolvedAdd precision, or a small reversible bet
Above zero but entirely below breakevenCreates revenue, not enough to cover costRework or cut

Your breakeven iROAS is 1 divided by your contribution margin, not 1.0x. At a 40% margin it is 2.5x, so a statistically significant 2.2x is still a loss. Significance tells you a channel did something; only breakeven tells you it paid. The full margin math is in what an iROAS confidence interval actually tells you.

Step 3: Choose your method from the table above: geo for channel or brand spend, holdout for campaign or audience-level questions.

Step 4: Design for a clean read. A matched or synthetic control for geo tests, one test per audience at a time, and a cooldown for longer conversion cycles. Google's geo-lift docs recommend the cooldown, and a campaign can sit in only one study at a time.

Step 5: Apply the rule and document what changed and why, so the program builds knowledge instead of a drawer of one-off decks.

Step 6: Move to the next question, and repeat.

How do you test a channel portfolio?

Compare channels instead of testing one in isolation. Suppose your contribution margin is 35%, so your breakeven iROAS is 2.9x. A holdout puts non-brand search at 3.4x and display retargeting at 1.6x. Retargeting reads positive and would pass a naive 1.0x check, but 1.6x sits well under 2.9x, so on the margin it loses money. Move that budget to search and re-test next quarter to confirm the new baseline.

That is reallocation on measured contribution, not on dashboards built to flatter every channel. In the 2024 IAB/PwC report, search and social each cleared $88B a year, and each reports itself a winner. Only your own test ranks them at your margin.

How does testing calibrate your media mix model?

An incrementality test is a causal reading at one point. A media mix model is an aggregate regression across your whole history, and left alone it tends to confuse demand with spend, since the two often move together. Feed the test results in as priors and each channel's coefficient is anchored to measured lift instead of correlation. That calibration is what separates an MMM that estimates causation from one that just fits a curve.

A calibrated model earns a careful forward look. It pulls seasonality, promotions, and price out of the media signal, so it does not credit ads for a holiday spike, then simulates how a budget split should perform and where the next dollar returns most. That projection carries error bars, holds only near spend levels you have run, and assumes conditions stay put. It is a forecast, not a look into the future, and it drifts the moment the calibrating tests go stale.

Incrementality testing vs A/B testing

Different questions. An A/B test compares two creatives or pages inside a channel and tells you which one wins. An incrementality test compares an exposed group against an unexposed one and tells you whether the channel added revenue at all. A winning A/B result says nothing about whether the channel earns its budget.

A/B testIncrementality test
Question it answersWhich ad or page performs better within a channelWhether the channel drives incremental revenue at all
What it comparesVariant A vs variant B, same audienceExposed vs unexposed, via geo or holdout
Tells youThe better creativeWhether to fund the channel
Reach for it whenOptimizing inside a channel you already trustDeciding whether a channel deserves budget

Frequently asked questions

How often should you run incrementality tests?

At least one test completing every month, rotated across channels. There is no fixed shelf life for a result: it decays as your creative, competition, and seasonality move, so keep the evidence about as fresh as the budget decisions it informs.

How long does an incrementality test stay valid?

There is no set expiry. The estimate is tied to the conditions inside the test window, and creative fatigue, competitor bidding, seasonality, and delivery all drift afterward. Re-test when those change, not on a fixed calendar.

Is the revenue lost during a holdout test wasted?

No. The withheld revenue is a research cost, small next to the spend the result governs. A holdout cannot prove a channel does nothing, but it can bound its lift below breakeven, which is enough to redirect the budget.

What is always-on incrementality testing?

A holdout or geo split that runs continuously and refreshes your estimate as conditions change. It has to be sized for power and read sequentially, and in return it flags a fading channel months before a periodic test would.

Should I run a geo experiment or a holdout test?

Geo for channel or brand spend and new channels, splitting regions with matched markets or a synthetic control. Holdout for campaign or audience-level questions with solid conversion coverage. Many programs rotate both.

Can a media mix model predict future performance?

It forecasts, within limits. A model calibrated with incrementality tests projects how a budget split should perform and where to reallocate, after separating seasonality and saturation from media. The projection carries error bars, holds only near spend levels you have run, and decays as the calibrating tests age.

Stella runs continuous incrementality and calibrated media mix modeling, so the numbers under your budget stay current. Book a demo.