How it works

The methodology, in full.

You don’t need to understand the math to use Stella. But you should be able to check our work.

No measurement method sees everything.

Each answers a different question. Stella uses them together because their blind spots are different.

Multi-touch attributionWhat did the customer touch?
Post-purchase surveysWhat does the customer remember?
Holdout experimentsWhat actually changed because of marketing?
Media mix modelingHow does performance change as spend changes?
Always-on measurementWhat’s changing right now?

None is the answer on its own.

Not every signal deserves the same weight.

The method matters. So does the quality of the evidence behind it. A strong holdout should carry more weight than a weak one, and a fresh, well-calibrated model more than a stale one. Stella evaluates the type, quality, recency, uncertainty, and agreement of the available evidence before using it.

How much each read counts toward the truth
Holdout experiment
anchor
Calibrated MMM
Automated daily study

Recent studies are weighted higher, so accuracy compounds as you keep testing.

Stella keeps learning

Every study makes the next answer better.

Always-on measurement turns your everyday spend changes into a steady stream of causal reads. Holdouts and MMM provide the stronger, structural evidence that calibrates what Stella believes. Attribution and surveys add the customer-level context. Every new signal updates the picture, and the decision it feeds updates with it.

Holdouts
Gold-standard geo experiments
MMM
Calibrated by your holdouts
Always-on
Live incremental factor, per channel
MTA
Weights every touchpoint
Ad platforms
Optimize for incremental margin
Holdouts
Gold-standard geo experiments
MMM
Calibrated by your holdouts
Always-on
Live incremental factor, per channel
MTA
Weights every touchpoint
Ad platforms
Optimize for incremental margin
↻ …and it loops, the more you run, the sharper the signal.
Validation

We don’t grade every method the same way.

A model proves itself differently than an experiment. So we show the checks that matter for each. A media mix model earns trust by predicting data it has never seen. A holdout earns trust from how closely the control tracked the test market before anything changed.

Both sets are shown in the product, on every study. An in-sample R² on its own is trivial to inflate, which is why it is never the headline here.

Media mix models

Predictive validation: does it predict data it never saw?

Out-of-sample R²0.88
Out-of-sample MAPE6.2%
VIF (collinearity)1.7
Forward and back testedyes
Holdout experiments

Experimental validation: did the design earn our trust?

Pre-period R² (control vs test)0.97
Pre-period MAPE2.4%
Then, the estimate itself
iROAS range1.8x to 2.4x
Total impact trendshown

Illustrative. The range is the result and its uncertainty, not a validation metric, which is exactly why it is shown separately from the design checks. No out-of-sample figure appears here, because a holdout does not have one.

Method by method

Know what each method can and can’t tell you.

Geo holdout experiments

Best forDid this marketing actually cause additional outcomes?
StrengthDirect causal evidence from a controlled comparison: a matched control region shows what would have happened anyway.
LimitationSlow (a clean read takes 20 to 28 days) and depends heavily on good test/control design.

Media mix modeling

Best forHow should performance change as spend moves, across every channel at once?
StrengthCovers channels you cannot click, and extends evidence across spend levels. Calibrated against your experiments.
LimitationObservational, so it rests on assumptions. We publish the out-of-sample numbers so you can judge the fit.

Always-on causal measurement

Best forWhat is incrementality doing right now, between formal tests?
StrengthEvery meaningful spend change becomes an automated before-and-after study, so the read stays current.
LimitationDirectional on its own, which is why it is weighted lowest and anchored by holdouts and MMM.

Multi-touch attribution

Best forWhat path did the customer actually take, at campaign and creative granularity?
StrengthThe only method that shows the real sequence of touches behind each sale.
LimitationA path is a sequence, not a counterfactual. It cannot say what your ads caused, or see channels with nothing to click.

Post-purchase surveys

Best forWhich channels do buyers themselves credit, including the ones no pixel sees?
StrengthCaptures TV, podcasts, word of mouth: the discovery channels invisible to tracking.
LimitationSelf-reported and subject to recall bias, so it is evidence, never proof.

So Stella weighs them together.

Five imperfect signals, weighed by the evidence behind them, resolved into one current answer: what caused growth, what still has room, and where the next dollar goes.

Attribution
Post-purchase surveys
Always-on causal
Media mix modeling
Holdout experiments
Stella
Weighs each signal by the strength of the evidence behind it.
One current answer
What caused growth. What still has room.
Where the next dollar goes

The method matters. So does the quality of the evidence behind it. Weights are illustrative: a weak study of any kind earns less than shown here, a strong one more.

FAQ

Definitions, plainly.

Why doesn’t Stella just use holdouts for everything?

Because holdouts are slow, and marketing is not. A clean geo holdout takes 20 to 28 days and answers one question about one channel at one spend level. Holdouts are the strongest evidence we have, which is exactly why we use them to calibrate everything else: the models and always-on measurement extend that rigor across every channel, every day, between tests.

Why can’t attribution tell me what caused a sale?

Attribution records the path a customer took: what they clicked and saw before buying. A path is a sequence, not a counterfactual. It cannot tell you whether the sale would have happened without those touches, and it cannot see channels with nothing to click, like TV or a podcast. That takes a controlled comparison, which is what holdouts are for.

What is incrementality testing?

Incrementality testing measures the revenue that only happened because of your advertising, the lift you would lose if you turned a channel off. Unlike attribution or ROAS, which count conversions that may have happened anyway, incrementality isolates true causal impact, usually with a geo holdout where some regions see ads and matched control regions do not.

What is a holdout study?

A holdout (or geo holdout) is a controlled experiment: you keep ads on in test regions and off in matched control regions, then compare. The difference is causal lift. Stella designs the test/control split, accounts for confounders, and runs an ensemble of models so the result is defensible, even to a skeptical CFO.

What is continuous causal optimization?

Every causal estimate starts decaying the moment you measure it: creative fatigues, CPMs rise, spend scales past what was tested. Continuous causal optimization keeps your incrementality numbers current with always-on measurement and calibration, so you act on what is true today, not on a test from last quarter.

What is a weighted synthetic control?

Instead of matching one control region to your test region, a weighted synthetic control blends many control markets into a synthetic twin that best reproduces your test market history. The closer that twin tracks before the test, the more trustworthy the lift estimate after it.

How do you decide which holdout result to trust?

We run your data through multiple models and present the one whose control best tracked your test market in the pre-period, before ads went off, not the one with the best-looking iROAS. You see the pre-period R² and MAPE, plus the iROAS range and statistical significance.

Is automated measurement as accurate as a holdout?

On its own an automated read is directional, which is why Stella weights it low. It becomes powerful when calibrated by your holdouts and MMMs: those anchor the truth, and the always-on signal inherits their accuracy between tests.

How does a survey answer become a number I can act on?

Answers map to canonical channels, so responses roll up into a share of discovery and the revenue behind it. Because only a portion of buyers answer, Stella projects those splits across your total revenue rather than pretending the surveyed orders are the whole business.

How is Stella different from legacy enterprise measurement vendors?

Two ways. First, transparency: we show out-of-sample MAPE, R², and VIF (the numbers that prove a model actually predicts), not just the in-sample R² that is trivial to fudge. Second, price and flexibility: full self-serve at $6,000/mo (about half a typical enterprise quote), a free tier for real holdout studies, or a fully managed engagement if you want us to run everything.

Bring us a study you’ve already run.

Run the same holdout through Stella and compare the result yourself.

Start freeor talk to a founder