Your MMM Is Lying About ROAS: The 2.5x Overstatement Nobody Calibrates
A Zalando researcher's simulation showed a standard media mix model reporting 10.61x ROAS for a channel whose true return was 4.20x — a 2.5x overstatement that better data cannot fix. Here is why the bias happens, what geo experiments can actually recover, and the three signals we extract before we move a single Q4 dollar.
A competently built media mix model can tell you a channel returns 10.61x when the real number is 4.20x — and no amount of better data fixes it. That's the finding of a simulation paper posted to arXiv on August 21 by Niklas Heusch, a researcher working from a zalando.de contact address. In a fully-known synthetic world, a standard MMM specified the way a competent practitioner would specify it — promotion dummy, observed price, three annual Fourier harmonics, the same seasonal controls Robyn, Meridian and pymc-marketing use — reported a 2.5x overstatement of paid search ROAS, and its 90% credible interval (6.56 to 14.36) didn't even contain the truth. The uncomfortable part is what happens next: hand the model the actual confounders — true promotion state, true price level, latent seasonality, quality drift, market sentiment — and it still says 8.41x. The bias is structural, not a data-quality problem. Budget decisions made on MMM alone are suspect, and with Q4 planning season here, that's a problem you need a process for, not a hope for.
Why the number travel makes it worse
The number that will travel is 10.61 against 4.20. It's a simulation, which is the only setting where you can know the true return at all — but that's exactly why it matters. Real practitioners never get to check their MMM against reality. They see "paid search: 10.6x ROAS" in the deck and allocate accordingly. The paper's core contribution is showing that the model was specified well: standard controls, standard tools, no strawman. The 2.5x gap is what competent modeling delivers.
The second number matters more for your process. When the same model was fed the true data-generating covariates (the "oracle" specification, which no practitioner possesses), it still returned 8.41x with an interval of 6.91 to 9.81. That bounds what better controls could ever achieve: roughly double the true return. It kills the comfortable story that richer data, cleaner logs, or a better seasonality prior will fix your MMM. The paper states it in one line: marketing budgets are not randomly assigned. Brands spend more ahead of expected demand, bidding systems raise spend when conversion rates rise — and high conversion rates usually reflect strong demand, not strong ads. The model attributes that correlation to causation, and no covariate you can hold constant absorbs an algorithm that responds to the random component of last week's performance.
The experiment you're already wasting
Here's where the paper turns practical. Geo experiments — go-dark tests split across DMAs — are supposed to be the calibration layer that fixes the MMM (Google shipped Meridian GeoX in May 2026 precisely to feed experiment results into the mix model). But the paper argues the standard handover throws away most of the information the experiment produced.
A well-designed geo test has three phases: a pre-test period establishing parallel trends, a test period where spend differs (often going dark in treatment), and a cooldown period where spend resumes and effects decay. The raw output is rich: daily time series of outcomes and spend for both groups across all three phases. Standard practice then collapses the whole thing into one number — total incremental revenue divided by total incremental spend — a point estimate that answers a narrow question about average return at one spending level.
Two signals die in that summation, and both are visible in the raw series:
| Signal | Where it shows up | What it recovers |
|---|---|---|
| Carryover (adstock decay) | How fast the outcome gap opens after go-dark and closes after cooldown | The decay rate that decides how long yesterday's spend keeps working |
| Diminishing returns | Tests at different spend levels | The saturation curve that decides where the next dollar stops earning |
| Effectiveness | The gap level itself once dynamics are separated | The coefficient that turns spend into revenue at all |
That's the shift the paper proposes: from "experiments for measurement" to "experiments for structural estimation." Instead of feeding Meridian a single ROAS point, you feed it the time series and let the model recover the three parameters every budget allocation depends on: persistence, speed of diminishing returns, and overall effectiveness.
What this means for your social budgets
We run paid media across Meta, TikTok, LinkedIn and Google for clients, and this paper changes how we read every aggregate model we see. Platform ROAS from Meta or LinkedIn is a different beast — it's in-platform signal, and we've written before about how engage-through attribution shifts and reporting breakdowns can silently distort it. But MMM sits on top of all of it, and if paid search ROAS is overstated ~2.5x in a known world, the social channels in the same model deserve the same skepticism — especially when the model has less data per channel to work with.
| Source | What it can tell you | What it can't |
|---|---|---|
| Platform ROAS (Meta/LinkedIn dashboards) | Which ad, audience or creative performs best in-platform | Whether the sale would have happened anyway |
| MMM with standard handover | Directional channel allocation at portfolio level | True ROAS — structurally biased, proven 2.5x overstatement |
| Geo test collapsed to one ROAS | A sanity-check point estimate | Carryover and saturation dynamics |
| Geo test fed as time series | Decay, response curve, effectiveness | Nothing — it's the closest thing to ground truth you can operationalize |
None of this means MMM is useless. It means MMM is a hypothesis generator, not a verdict. When we calibrate a client's model before Q4, we don't ask "what does the model say?" — we ask "which channels are we planning to shift, and what test evidence do we have for those shifts?" If the answer is no test evidence, the shift gets an experiment budget and a decision rule before it gets a reallocation.
The three-signal handover we run before Q4
If you have a geo test coming up — or you're about to trust one from last quarter — this is the minimum extraction we'd do:
1. Extract the decay, not the average. Look at the go-dark transition: does the outcome gap open in a week or across six? That transition rate is the adstock decay parameter your MMM is guessing. A test that collapses to one ROAS figure discards it. We've seen clients' models assume 3-week carryover when the true decay was under a week — that single miss misallocates more budget than any creative tweak that quarter.
2. Trace the curve, not the point. If the test ran two or more spend levels, the difference between them traces the saturation curve. Doubling spend and doubling the gap means proportional returns; less than double means saturation — and the shape tells you where the next dollar actually earns.
3. Keep the point estimate as a sanity check, never as the decision. Summed incremental revenue over summed incremental spend is still a fine headline for a deck. It's a terrible input for a model update. Meridian's GeoX and similar tools can take the richer feed — give them the signal they're designed for.
4. Agree on the decision rule before the test starts. If the experiment shows incremental ROAS below your blended target, what happens — pause, trim 20%, reallocate to creative? If you haven't written that down before the data lands, you'll rationalize whatever the dashboard shows. This is the same discipline we apply to channel incrementality tests: the rule comes first, or the test is just a spending exercise.
5. Reconcile the model against reality quarterly, not annually. The Q4 budgeting cycle is the obvious checkpoint, but the paper's bias doesn't take holidays: bidding algorithms drift, promo calendars change, baseline demand moves. A once-a-year calibration means your model is wrong for eleven months, and you just don't know which eleven.
The practical takeaway for a marketing team heading into Q4: treat every MMM number over 2x your blended target as a flag, not a fact. Run or refresh the geo tests on the channels you actually plan to shift. Extract the three signals. Then move money with evidence — not with a model that, in a world where the truth was known, was wrong by a factor of 2.5 and confident about it.
Sources: Heusch, N. — "Structural Estimation of Marketing Mix Model Parameters from Geo-Experiments," arXiv:2608.21128 (Aug 21, 2026); PPC Land coverage (Aug 31, 2026); Google Meridian GeoX announcement (May 2026).
Frequently Asked Questions
Why do media mix models overstate ROAS?
Because marketing budgets are not randomly assigned. When demand rises, brands and bidding algorithms spend more — the model sees correlated spend and sales and credits the ads with sales that would have happened anyway. In the Heusch simulation, even an oracle model given the true confounders still returned 8.41x against a true 4.20x, so better data alone cannot close the gap.
What is the difference between a standard geo test handover and structural estimation?
Standard practice collapses the test into one point estimate (total incremental revenue divided by total incremental spend). Structural estimation instead feeds the full time series into the model, which lets it recover the adstock decay rate, the saturation curve, and the effectiveness coefficient separately — the three parameters every budget decision depends on.
Should I move budget based on my MMM before Q4?
No, not on the model alone. Use the MMM as a hypothesis generator, then validate the channels you plan to shift with a refreshed geo or holdout test, extracting the carryover, saturation, and effectiveness signals before reallocating. A single ROAS number should never trigger a reallocation on its own.



