When Should a Company Start Incrementality Testing?

Start incrementality testing when your design can reliably detect a lift small enough to matter for the decision you are making, and when the test can finish before your account changes out from under it. For most DTC brands, a useful first screen is roughly 100 or more daily orders in a single country. Monthly spend on one channel, very roughly $75,000 to $100,000 on Meta or Google, tends to track the same thing. Treat both as rules of thumb, not statistical cutoffs. Below that range, the honest answer is usually to wait and grow first.
Most guides answer this question with a spend number and stop. Spend is easy to look up, so it makes a tidy threshold. It is also the wrong variable to lead with.
What decides whether a test works is not how much you spend. It is whether a holdout can gather enough orders to see an effect through the noise.
Holdouts are widely misunderstood, so it is worth being precise. A holdout compares two sets of regions whose sales normally move together, matched by location selection analysis. In a geo holdout you switch something on in the test markets, a new channel or a new ad, and leave the matched markets running as usual. An inverse geo holdout is the reverse: you turn a channel off in the test markets while the rest stay live. Under a valid design, where the markets stay comparable and nothing else pulls them apart, the gap that opens up between the two groups is your estimate of the lift the ads caused. Get the market matching, the timing, or the isolation wrong and that gap can be something else, which is why design quality matters as much as volume.
Standard platform attribution is counting, not causation. The conversions on your dashboard include purchases that would have happened without the ad. Last-click attribution hands all the credit to whatever the shopper touched last, which is often a channel just collecting a sale that was already going to happen. Incrementality testing is how you find out what your ads are really adding.
What actually decides if you're ready to test?
Conversion volume and noise, not spend. What you can measure comes down to the smallest lift your test can reliably catch, its minimum detectable effect. That depends on several things at once: how much your daily sales bounce around, how closely your test and control markets track each other, how long you run, and how big a lift you are hoping to find. Order volume is one of the biggest levers on it, which is why a brand doing 100 orders a day can usually resolve a normal-sized lift while a brand doing a handful cannot inside a sensible window. Spend matters only because it tends to track order volume.
Here is the mechanic in plain terms. Statistical significance asks whether the gap you measured is large relative to the noise the test would expect if your ads had no effect at all. To get there, a test compares matched groups of markets, and each group needs enough steady order flow that ordinary swings do not swamp the signal.
When volume is low, the two groups swing around from week to week for reasons that have nothing to do with your ads. A small true lift gets buried inside that ordinary variation, and the test comes back inconclusive.
When volume is high, those swings shrink relative to the signal, so a smaller lift becomes visible. Order count is not the only thing that matters, but it is one of the main levers on how small an effect you can reliably catch.
This is why we start with a look at daily order volume before anything else. The number tells you what is measurable before you spend a day testing.
What are the thresholds for incrementality testing?
Two numbers to sanity-check. The one that matters is about 100 or more daily orders in a single country. The rough proxy is somewhere around $75,000 to $100,000 a month on one channel like Google or Meta. Enough daily orders usually gives a holdout the power to resolve a real lift instead of coming back inconclusive. When spend and orders disagree, trust the orders. Both are screening heuristics from the accounts we see, not formulas.
Why orders matter more than spend: two brands can spend the same and be in completely different shape to test. A brand with a low average order value pushes far more orders through that budget than a high-AOV brand does, and it is the order count, not the dollars, that gives a holdout its power. A high-AOV brand can clear $100,000 a month and still sit below the volume a clean test needs. That is the whole reason we lead with orders.
Keep it to a single country. Testing across borders drags in currency, seasonality, and shipping differences that widen the noise you are trying to shrink.
Treat both numbers as starting lines, not magic cutoffs. Your real floor depends on how noisy your daily sales are and how cleanly your markets can be matched. A steady, predictable order flow, well matched, resolves a lift sooner than a spiky, hard-to-match one at the same volume.
| Daily orders (one country) | What a holdout can do | Verdict |
|---|---|---|
| Well under 100 (a few dozen a week or fewer) | Usually too little signal to separate a real lift from weekly swings, and long runs bring their own problems. | Wait |
| Approaching 100 | A first, simple test can work on your single largest channel. Expect a longer window, longer still if your AOV is high. | Edge: test one clear question |
| 100 or more | Usually enough signal for a well-matched holdout to catch a normal-sized lift within a typical window. | Test-ready |
How long should an incrementality test run?
Long enough for the holdout to capture the sales your ads caused, which depends on your ad-to-purchase timeframe. If shoppers take three weeks to decide, a one-week test misses most of the effect and reads low. High average order value usually means a longer consideration window, so high-AOV brands need a longer holdout. Across 225 DTC brands, Stella tests ran 20 to 59 days, with a median of 33.
That 225-test figure is a self-selected group of brands that chose to run a test with us, not a market average, so read it as a real spread rather than a promise. What it shows is useful anyway: most tests resolve in about a month, and almost none in a single week.
The reason duration moves is the gap between seeing an ad and buying. A $30 impulse product converts fast, so the sales show up inside a short window. A $600 considered purchase can take weeks, and a holdout that ends before those delayed orders land will undercount the lift and make a good channel look dead.
So the length of your test is not a fixed setting. It follows your own purchase cycle. Match the window to how long your customers actually take to buy, then give it room. You can see the full duration and iROAS spread in our benchmarks.
When is incrementality testing a waste?
When you are too small to fill a holdout. If paid drives only a few dozen orders a week, you usually cannot clear the noise, and running longer rarely rescues it because it brings problems of its own. The longer a test runs, the harder it is to hold the treatment and the business steady. Creative rotation, promos, price changes, competitive moves and seasonality can all shift what the test is even measuring.
That does not mean you skip measurement. A holdout is one of three ways Stella measures, and it is the one that needs scale. Below the volume a clean holdout needs, a post-purchase survey that asks buyers where they actually heard about you, and multi-touch attribution that weighs every touch instead of only the last click, can still point you in the right direction. They answer different questions than a lift test and are not a causal substitute for one, but they beat running on last-click alone while you grow into a test. Stella runs both, post-purchase surveys and multi-touch attribution.
This is the trap small accounts fall into. Too few conversions to shorten the test, too much creative churn to lengthen it. Squeezed from both ends, the test cannot win, and a month of held-out revenue buys you a result you cannot trust.
Most testing vendors will not tell you this, because a test is what they sell. We will. There is a real point where a test is the dumbest thing you could do, and it is worth knowing you are past it before you hold out a dollar of revenue.
How do you run your first test?
Start with one channel and an inverse geo holdout: turn ads off in matched regions, hold every budget flat, and measure against total store revenue, not platform-reported conversions. Freeze spend for the whole window. The gap between your incremental ROAS and your platform-reported ROAS is the number you are after.
Measure against total revenue because that is what you actually care about. Platform-reported conversions are counting, not causation. Incremental ROAS is the revenue your ads caused divided by what you spent, and it usually differs from the figure on the dashboard, often lower, sometimes higher. The direction is not the point. The point is that attribution and causality are answering different questions.
How big that gap is, and which way it runs, is what a test tells you. In a set of 15 large field experiments on Facebook, observational methods missed the randomized result badly, overstating the true ad effect by roughly three times in half of them, per Gordon et al. (2019) in Marketing Science. The lesson is not that every dashboard is exactly 3x too high. It is that you cannot tell how far off yours is, or in which direction, without a test.
Real ad-driven return is a long way from zero. Across 225 DTC brands that tested with us, a self-selected group that chose to test and passed our pre-test screening, not a market average, the middle half of tests landed between 1.36x and 3.24x incremental ROAS, with a median of 2.31x. That is the distribution across those 225 tests, not a promise about your account.
Freeze budgets while the test runs. Change spend mid-test and you can no longer tell the effect of the change apart from the lift you were trying to measure, and the read is gone.
Ready to find out what your ads are actually adding? Book a demo.
Frequently asked questions
How many daily orders do you need to run an incrementality test?
Around 100 or more daily orders in a single country is a workable starting point for most e-commerce brands. That gives a holdout enough conversions to detect lift through normal weekly variation. You can test with fewer, but expect a longer window and a higher chance the result comes back inconclusive.
Can a small e-commerce brand run incrementality testing?
It can, but often should not yet. If paid drives only a few dozen orders a week, a holdout cannot gather enough conversions to separate real lift from ordinary swings, and extending the test past a month runs into creative turnover. Growing order volume first usually beats testing early.
Does a high-AOV brand need a longer test?
Usually yes. High average order value tends to come with a longer consideration window, so sales land days or weeks after someone sees the ad. A short holdout ends before those delayed purchases show up, so high-AOV brands hold their test regions out longer to capture the full effect.
Is monthly ad spend or daily sales the better readiness signal?
Daily sales. Spend is a rough proxy that matters only because it tends to track order volume. Two brands spending the same amount can have very different order counts depending on price point, and it is the order count, more than the spend, that shapes whether a holdout can resolve a real lift.
How long does an incrementality test take to reach significance?
It depends on your order volume and your ad-to-purchase timeframe. Across 225 DTC brands that ran a Stella test, durations ranged from 20 to 59 days, with a median of 33. Higher volume and shorter purchase cycles resolve faster. Low volume and long consideration windows take longer.