How Do You Turn Measurement Results Into a Budget Decision?

You turn an incrementality or MMM result into a budget decision by working backward from the decision, not forward from the number. Start with the budget question, like whether to add $50K to Meta. Confirm the test or model can actually answer it. Then write a recommendation that states the direction, the size, the uncertainty, the evidence type, and the date you'll check whether it worked.
Picture the quarterly review. Slide 14 says Meta is 2.9x iROAS. Somebody asks "so we scale Meta?" and the room goes quiet. Nobody can say by how much, or how sure to be. The math was fine. Nobody wrote the path from the number to the move.
Why do most measurement results never change the budget?
Because the output stops at a number. A lift estimate or an iROAS table describes what happened. It doesn't say what to move or how far. When nobody writes that part down, the budget defaults to last year's plan with a nicer deck attached.
In a Google-sponsored Harvard Business Review Analytic Services survey of 547 marketers familiar with their company's MMM (September and October 2025), 87% said using MMM for data-driven insights is important. Only 28% said they're very effective at turning those insights into timely action.
Holdouts have the same problem. A number with no recommendation is safe for whoever delivered it, because nobody can grade it. "Raise Meta 15%" can be graded. That's why you want it in writing.
Can a holdout answer the decision you need?
Check before you commit. A location selection analysis has two jobs before anything else. Show that the spend change is big enough to answer the decision, and that the channel spends that much today. Then show that the available markets can support a credible comparison. Fail either one and the test can't settle the decision.
The first job is money. Detecting a lift depends on the design, revenue noise, how many markets you have and how different they are, test length, spillover, and pre-period fit. Spend is the knob you control. So the analysis finds the minimum investment: the least spend you can turn off (or add, in a scale-up test) and still reliably detect the smallest lift that would change your decision. That takes an assumed effect size, a design assumption, not a test result.
Say you'd need to hold out $60K and the channel only spends $35K in those markets (illustrative numbers). You can't hold out money you aren't spending. Add markets, run longer, flip to a scale-up test, or ask a different question. Don't run it anyway. An underpowered test comes back too uncertain to settle anything, and teams read that as "the channel doesn't work." Our advanced incrementality guide covers the minimum detectable effect.
The second job is the markets. The analysis looks for control markets that predict the test markets with low, stable error across the pre-period. High correlation helps, but it's a supporting check, not the bar. Good designs also run a placebo on the pre-period (does the method find a "lift" where nothing changed?) and check for spillover across market borders. More in our practitioner's guide to weighted synthetic control.
The minimum investment is your first budget decision. A holdout saves the media you turn off and loses the contribution margin on the sales that go away. Pull $50K of spend, lose $100K of revenue at a 40% margin, and you give up $40K of contribution to save $50K. You come out $10K ahead, because the channel was below breakeven. So a holdout costs roughly the profit the channel was making there. If that price beats what the decision is worth, skip the test.
What should a written budget recommendation include?
Six parts. The direction, including where the money comes from. The size of the move. The interval, measured against your breakeven multiple. Whether the channel was cleanly separated from the others. Whether the evidence is experimental, modeled, or both. And the check date, with the result that would reverse the call. Miss one and you have an opinion with a slide attached.
| Part | What it looks like | What it looks like when it's missing |
|---|---|---|
| Direction | Raise Meta, funded by cutting display. Hold branded search. | "Meta is performing well" |
| Size | +15% next month | "Lean into Meta" |
| Interval vs. breakeven | 2.6x to 3.3x against a 2.5x breakeven | "2.9x iROAS" with no range |
| Identification (MMM) | Meta's estimate held up when the model was rebuilt without YouTube's tangle | Clean per-channel iROAS for channels that always ramp together |
| Evidence type | Geo holdout plus an MMM calibrated to it | Unstated, so you assume the strongest |
| Check date and reversal | Backtest monthly. Revenue below the predicted range two months running triggers a retest. | Nothing, so the call can never be wrong |
Interval vs. breakeven. iROAS is a revenue multiple, not a profit one. Breakeven is 1 divided by your contribution margin (after shipping, returns, and fees), so a 40% margin means 2.5x. A 2.9x result with a 90% interval of 2.6x to 3.3x clears breakeven across the whole range. The same 2.9x at 1.8x to 4.0x doesn't. For scaling, compare breakeven to the return on the next dollars, which isn't always what the test reported.
What "interval" means depends on the method: a confidence interval from a frequentist test, a credible interval from a Bayesian MMM, or a randomization-based interval from a geo test.
Identification. Can the model tell one channel's effect from another's? If Meta and YouTube rose and fell together all year, it can't split credit cleanly. VIF flags that risk, but no cutoff makes a channel iROAS safe. The better test is whether the call survives the model being rebuilt a slightly different way. More in the multicollinearity post.
Evidence type. A holdout changes spend on purpose. An MMM recovers effects from variation that happened on its own, which takes stronger assumptions. Both are causal measurement, unlike last-click, but they don't carry equal weight. If the MMM was calibrated with the same holdout, their agreement is one piece of evidence counted twice. Even uncalibrated, it runs on the same revenue history, so agreement is corroboration, not independent replication.
Check date and reversal. In Stella's language, the backtest is the live check after the money moves. (A forward test is different: the model hides a stretch of data it already has, predicts it, and gets compared to what actually happened.) Each month, compare actual revenue to what the pre-decision model specification predicts for the new mix, given the prices, promos, and spend that actually happened. That's monitoring, not validation. Validation comes from a holdout or a retest. So write the trigger down first: two months below the predicted range, with no obvious explanation, and the moved channel gets a holdout before any more scaling.
How big should the budget move be?
Size it to the downside. Take the low end of the interval for the dollars you'd add, work out what you'd lose on each one if that's the truth, and set the step so that loss stays inside a number you've agreed to absorb. Then check the money wouldn't earn more somewhere else.
That number is a loss budget: how much contribution you'll risk this cycle if the evidence lands on the low end. Finance and marketing set it before the result comes in. Moves you can reverse next month (paid social, search) can carry a bigger one than moves you can't (a TV upfront).
The math, with illustrative numbers. Margin is 40%, so breakeven is 2.5x. The return on the next $50K into Meta is 2.9x, with a 90% interval of 2.1x to 3.7x. Each extra dollar returns 0.4 times the iROAS in contribution, minus the dollar:
- At 2.9x, each extra dollar earns $0.16.
- At the 2.1x low end, each extra dollar loses $0.16.
With an $8K loss budget, the most you can add is $8K divided by $0.16, or $50K. On $250K a month, that's a 20% step. At the point estimate you gain about $8K. At the low end you lose the $8K you agreed to risk. The truth can land below that bound, so the loss budget is a plan, not a guarantee.
That interval straddles breakeven and still earned a real move, because being wrong is survivable. The math treats 2.9x as constant across the step, which is fair for small moves. For bigger ones, evaluate along the response curve.
Breakeven is only the floor. If search can absorb the next $50K at 4.2x, Meta's 2.9x is the worse use of the money. The loss budget tells you how much risk to take. Comparing marginal returns tells you where.
The 90% level is a choice. Stella uses 90% on live tests because monthly moves are reversible. The loss budget sizes your exposure to that trade at the downside bound you picked in advance. Irreversible calls deserve 95%, which is wider, so the same loss budget buys a smaller step.
| Evidence situation | 90% interval vs. breakeven | Starting guidance |
|---|---|---|
| Holdout that passed feasibility and pre-period checks, plus an MMM not calibrated to it that reaches a consistent conclusion | Low end clears breakeven | Step up, capped near the spend range you tested. Backtest monthly. |
| Holdout only | Low end clears breakeven | Step up near the tested range. Run a scale-up test before going far past it. |
| MMM only, channel call stable across rebuilds | Low end clears breakeven | Smaller step. Schedule a holdout. |
| MMM only, channel tangled with another | Any | Hold that pair. Test to separate them. |
| Any evidence | Straddles breakeven | A step sized to the loss budget, or a hold, plus a retest |
| Decision-grade evidence that passed identification and diagnostics | High end below breakeven | Cut in steps. Watch the backtest for what else moves. |
| Any evidence | Clears breakeven, but another channel's next dollars clearly earn more | Move the money to the better channel, sized by the same loss budget |
A holdout tells you about the spend change you ran. Cut Meta by $20K a week and you learn what that $20K was doing, not what $20K more would do. If the estimate came from cutting spend, treat it as the optimistic case for scaling above today's level, unless you have evidence the response curve stays flat. More in average vs. marginal iROAS.
Hold is a decision. When the interval straddles breakeven and the upside is small, hold and retest. Cutting a channel because the point estimate looked soft is how brands kill channels that were working.
How do I get a real budget recommendation from my vendor?
Ask the ninth question from our tear-out card: what budget change do you recommend, and what uncertainty sits around that move? A strong answer names channels, amounts, the uncertainty, and a condition. Raise Meta 15%, funded from display, because the downside stays inside the loss budget. Hold branded search. Retest YouTube. A weak answer repeats the iROAS.
It comes last on purpose (vetting a measurement consultancy covers the other eight). A recommendation built on a model that can't predict, or can't separate channels, isn't worth sizing.
For holdouts, ask before you sign: did you check feasibility, and can I see it? A strong answer shows the minimum holdout investment against your spend and how well the chosen markets predicted each other. A weak answer is "we'll pick some markets and see."
Red flags:
- The output ends at an iROAS number, with no move attached
- A move with a size but no interval
- An interval measured against 1.0x instead of your breakeven
- A move into one channel with no word on where the money came from
- "Keep testing" with no test, trigger, or date
FAQ
What's the difference between a measurement output and a budget recommendation?
An output describes the past: lift, contribution, iROAS, response curves. A recommendation commits to a future move with a size, a confidence level, and a date to check it. The first gets presented. The second gets acted on, and graded.
Should I move budget on MMM alone, without a holdout?
You can, in smaller steps. MMM takes stronger assumptions than a holdout that changes spend on purpose. When a large move rests on modeled evidence alone, run a holdout on that channel first, or calibrate the model with one.
How often should I check a budget recommendation after it ships?
Monthly. Compare actual revenue to what the pre-decision model predicted for the new mix. A miss flags that something changed. A holdout on the moved channel tells you whether the move itself was wrong.
What does a holdout feasibility check look at?
Whether the spend change needed to detect a decision-relevant lift fits inside what the channel spends now. And whether the candidate markets predict each other with low, stable error.
Is "keep testing" a valid recommendation?
Yes, if it names the test, the result that would trigger a move, and the timeline. Without those, it's a way to avoid making one.
Every Stella result ships with the recommendation written out: the move, the size, the interval against your breakeven, the evidence, and the check date. Every holdout starts with the feasibility check. Book a demo to see one.