How it works

The methodology, in full.

You do not need any of this to use Stella. It is here because the number you take into a budget meeting should be one you can defend, and that means being able to check our work.

Validation

A model and an experiment are not validated the same way.

Vendors tend to quote one set of numbers for everything. That is a tell. A media mix model earns trust by predicting data it has never seen. A geo holdout has no out-of-sample metric at all: it earns trust from how closely the control tracked the test market before anything changed, and from how tight the resulting range is.

Both sets are shown in the product, on every study. An in-sample R² on its own is trivial to inflate, which is why it is never the headline here.

Media mix models

Validated out of sample.

Out-of-sample R²0.88
Out-of-sample MAPE6.2%
VIF (collinearity)1.7
Forward and back testedyes
Holdout experiments

Validated on the pre-period, then on the result.

Pre-period R² (control vs test)0.97
Pre-period MAPE2.4%
iROAS range1.8x to 2.4x
Total impact trendshown

Illustrative. No out-of-sample figure appears here, because a holdout does not have one.

Calibration

Each method calibrates the next.

Holdouts are the most rigorous evidence and the slowest to produce. Models are fast and only as good as what anchors them. So the holdouts calibrate the media mix models, both calibrate the always-on signal, and the incremental factor flows back into attribution and out to the ad platforms. The more you run, the sharper the signal.

Holdouts
Gold-standard geo experiments
MMM
Calibrated by your holdouts
Always-on
Live incremental factor, per channel
MTA
Weights every touchpoint
Ad platforms
Optimize for incremental margin
Holdouts
Gold-standard geo experiments
MMM
Calibrated by your holdouts
Always-on
Live incremental factor, per channel
MTA
Weights every touchpoint
Ad platforms
Optimize for incremental margin
↻ …and it loops, the more you run, the sharper the signal.
Method by method

What each one is good for, and where it falls down.

No single method is complete. The honest version is that each has a blind spot the others cover, which is the whole reason Stella runs them together.

Geo holdout experiments
The strongest causal evidence available, because a matched control region tells you what would have happened anyway. Slow: a clean read takes 20 to 28 days, and a badly matched control ruins it before any model runs.
Media mix modeling
Covers every channel at once, including the ones you cannot click, and extends evidence across spend levels. Observational, so it depends on assumptions; we publish the out-of-sample numbers and calibrate against your experiments.
Always-on causal measurement
Turns each meaningful spend change into an automated 15-day before-and-after study, so you see incrementality shift between formal tests. Directional on its own, which is why it is weighted lowest and anchored by the other two.
Multi-touch attribution
The only method that shows the actual path a customer took, at campaign and creative granularity. A path is a sequence, not a counterfactual, so it cannot tell you what your ads caused.
Post-purchase surveys
Captures the channels no pixel can see: podcasts, TV, a friend at dinner. Self-reported and subject to recall bias, so it is evidence, never proof.

So none of them wins. They get weighed.

Strength of evidence is a property of the study, not the method. A well-run experiment and a well-fit model both earn their weight; a weak one of either earns less. Stella weighs each signal on that basis and resolves them into one answer.

Attribution
Post-purchase surveys
Always-on causal
Media mix modeling
Holdout experiments
Stella
Weighs each signal by the strength of the evidence behind it.
One view of performance
Continuously updated as new evidence arrives.
Where the next dollar goes

Strength of evidence is a property of the study, not the method.

FAQ

Definitions, plainly.

What is incrementality testing?

Incrementality testing measures the revenue that only happened because of your advertising, the lift you would lose if you turned a channel off. Unlike attribution or ROAS, which count conversions that may have happened anyway, incrementality isolates true causal impact, usually with a geo holdout where some regions see ads and matched control regions do not.

What is a holdout study?

A holdout (or geo holdout) is a controlled experiment: you keep ads on in test regions and off in matched control regions, then compare. The difference is causal lift. Stella designs the test/control split, accounts for confounders, and runs an ensemble of models so the result is defensible, even to a skeptical CFO.

What is continuous causal optimization?

Every causal estimate starts decaying the moment you measure it: creative fatigues, CPMs rise, spend scales past what was tested. Continuous causal optimization keeps your incrementality numbers current with always-on measurement and calibration, so you act on what is true today, not on a test from last quarter.

What is a weighted synthetic control?

Instead of matching one control region to your test region, a weighted synthetic control blends many control markets into a synthetic twin that best reproduces your test market history. The closer that twin tracks before the test, the more trustworthy the lift estimate after it.

How do you decide which holdout result to trust?

We run your data through multiple models and present the one whose control best tracked your test market in the pre-period, before ads went off, not the one with the best-looking iROAS. You see the pre-period R² and MAPE, plus the iROAS range and statistical significance.

Is automated measurement as accurate as a holdout?

On its own an automated read is directional, which is why Stella weights it low. It becomes powerful when calibrated by your holdouts and MMMs: those anchor the truth, and the always-on signal inherits their accuracy between tests.

How does a survey answer become a number I can act on?

Answers map to canonical channels, so responses roll up into a share of discovery and the revenue behind it. Because only a portion of buyers answer, Stella projects those splits across your total revenue rather than pretending the surveyed orders are the whole business.

How is Stella different from legacy enterprise measurement vendors?

Two ways. First, transparency: we show out-of-sample MAPE, R², and VIF (the numbers that prove a model actually predicts), not just the in-sample R² that is trivial to fudge. Second, price and flexibility: full self-serve at $6,000/mo (about half a typical enterprise quote), a free tier for real holdout studies, or a fully managed engagement if you want us to run everything.

Bring us a study you have already run.

Put a holdout or a model you ran elsewhere through Stella's free tier and compare the results yourself. That is a faster way to judge the methodology than reading about it.

Start freeor talk to a founder