The methodology, in full.
You don’t need to understand the math to use Stella. But you should be able to check our work.
No measurement method sees everything.
Each answers a different question. Stella uses them together because their blind spots are different.
None is the answer on its own.
Not every signal deserves the same weight.
The method matters. So does the quality of the evidence behind it. A strong holdout should carry more weight than a weak one, and a fresh, well-calibrated model more than a stale one. Stella evaluates the type, quality, recency, uncertainty, and agreement of the available evidence before using it.
Recent studies are weighted higher, so accuracy compounds as you keep testing.
Every study makes the next answer better.
Always-on measurement turns your everyday spend changes into a steady stream of causal reads. Holdouts and MMM provide the stronger, structural evidence that calibrates what Stella believes. Attribution and surveys add the customer-level context. Every new signal updates the picture, and the decision it feeds updates with it.
We don’t grade every method the same way.
A model proves itself differently than an experiment. So we show the checks that matter for each. A media mix model earns trust by predicting data it has never seen. A holdout earns trust from how closely the control tracked the test market before anything changed.
Both sets are shown in the product, on every study. An in-sample R² on its own is trivial to inflate, which is why it is never the headline here.
Predictive validation: does it predict data it never saw?
Experimental validation: did the design earn our trust?
Illustrative. The range is the result and its uncertainty, not a validation metric, which is exactly why it is shown separately from the design checks. No out-of-sample figure appears here, because a holdout does not have one.
Know what each method can and can’t tell you.
Geo holdout experiments
Media mix modeling
Always-on causal measurement
Multi-touch attribution
Post-purchase surveys
So Stella weighs them together.
Five imperfect signals, weighed by the evidence behind them, resolved into one current answer: what caused growth, what still has room, and where the next dollar goes.
The method matters. So does the quality of the evidence behind it. Weights are illustrative: a weak study of any kind earns less than shown here, a strong one more.
Definitions, plainly.
Why doesn’t Stella just use holdouts for everything?
Because holdouts are slow, and marketing is not. A clean geo holdout takes 20 to 28 days and answers one question about one channel at one spend level. Holdouts are the strongest evidence we have, which is exactly why we use them to calibrate everything else: the models and always-on measurement extend that rigor across every channel, every day, between tests.
Why can’t attribution tell me what caused a sale?
Attribution records the path a customer took: what they clicked and saw before buying. A path is a sequence, not a counterfactual. It cannot tell you whether the sale would have happened without those touches, and it cannot see channels with nothing to click, like TV or a podcast. That takes a controlled comparison, which is what holdouts are for.
What is incrementality testing?
Incrementality testing measures the revenue that only happened because of your advertising, the lift you would lose if you turned a channel off. Unlike attribution or ROAS, which count conversions that may have happened anyway, incrementality isolates true causal impact, usually with a geo holdout where some regions see ads and matched control regions do not.
What is a holdout study?
A holdout (or geo holdout) is a controlled experiment: you keep ads on in test regions and off in matched control regions, then compare. The difference is causal lift. Stella designs the test/control split, accounts for confounders, and runs an ensemble of models so the result is defensible, even to a skeptical CFO.
What is continuous causal optimization?
Every causal estimate starts decaying the moment you measure it: creative fatigues, CPMs rise, spend scales past what was tested. Continuous causal optimization keeps your incrementality numbers current with always-on measurement and calibration, so you act on what is true today, not on a test from last quarter.
What is a weighted synthetic control?
Instead of matching one control region to your test region, a weighted synthetic control blends many control markets into a synthetic twin that best reproduces your test market history. The closer that twin tracks before the test, the more trustworthy the lift estimate after it.
How do you decide which holdout result to trust?
We run your data through multiple models and present the one whose control best tracked your test market in the pre-period, before ads went off, not the one with the best-looking iROAS. You see the pre-period R² and MAPE, plus the iROAS range and statistical significance.
Is automated measurement as accurate as a holdout?
On its own an automated read is directional, which is why Stella weights it low. It becomes powerful when calibrated by your holdouts and MMMs: those anchor the truth, and the always-on signal inherits their accuracy between tests.
How does a survey answer become a number I can act on?
Answers map to canonical channels, so responses roll up into a share of discovery and the revenue behind it. Because only a portion of buyers answer, Stella projects those splits across your total revenue rather than pretending the surveyed orders are the whole business.
How is Stella different from legacy enterprise measurement vendors?
Two ways. First, transparency: we show out-of-sample MAPE, R², and VIF (the numbers that prove a model actually predicts), not just the in-sample R² that is trivial to fudge. Second, price and flexibility: full self-serve at $6,000/mo (about half a typical enterprise quote), a free tier for real holdout studies, or a fully managed engagement if you want us to run everything.
Bring us a study you’ve already run.
Run the same holdout through Stella and compare the result yourself.