Retail Media
RGM° · Training
Retail Media Incrementality
Why reported ROAS lies. The methodologies, test designs, and decision rules that separate real lift from cannibalization.
Why retail media ROAS lies
Sponsored search on Amazon, Walmart, or Target reports 8×, 12×, or 20× ROAS routinely. Brand teams celebrate, scale spend, and the platforms report even better ROAS. Then quarterly sales review shows total category growth is flat or even declining despite ad spend tripling. What happened?
The answer is structural. Platform-reported ROAS is last-click attribution to shoppers already at the retailer searching for relevant products. A large share of those shoppers would have purchased anyway — especially on branded searches. The ad gets credit for purchases it didn't cause.
Industry research (Amazon's own MMM studies, Catalina case studies, Nielsen iROAS benchmarks) suggests that incremental ROAS (iROAS) is often 30–70% of reported ROAS. For branded sponsored search, iROAS can be as low as 10–20% of reported. For category prospecting and new-to-brand, iROAS can exceed reported because of halo effects on related SKUs.
This is not unique to retail media — the same dynamic applies to branded paid search on Google — but the scale of cannibalization is larger in retail media because the user's purchase intent is so high at the moment of ad exposure.
The three types of incrementality testing
| Type | How it works | Strengths | Limits |
| Geo holdout | Pause or reduce ads in selected geos; compare sales in test vs control geos | Works without user-level data; controllable; respects privacy | Requires sufficient geo separation; can't isolate audience-level effects |
| User-level holdout (ghost ads / PSA) | Random users in matched audience see ad; matched control sees public-service ad or no ad; compare conversions | Clean causal inference at user level; tight statistical control | Requires platform support; small audience sizes need long test duration |
| Media mix model (MMM) | Statistical model fitted to 2+ years of weekly sales, spend, and external factors; isolates channel contribution | Cross-channel attribution; accounts for adstock, saturation, halo | Slow (rebuild quarterly), expensive, requires data discipline |
Geo holdouts: the workhorse methodology
Geo holdouts are the most common incrementality methodology in retail media. The mechanics:
- Select test cells. Group DMAs or zip codes into 2–4 matched cells based on historical sales, geography, and demographics. Tools like Google's Geo Experiments framework or third-party tools (Haus, Recast, Lift) help.
- Assign treatment. Cells get treatment (ad on / off / spend levels). Common designs: A/B (50/50 split), A/A/B (33/33/33), or matched-pair (one cell heavy, matched cell light, third cell control).
- Run for sufficient duration. 6–12 weeks typical. Daily-purchase categories can be shorter; quarterly-purchase categories need longer.
- Measure lift. Compare sales (from retailer point-of-sale or syndicated data) across cells. Apply statistical test (t-test, regression, or causal-inference framework).
- Translate to iROAS. Lift / spend = iROAS. Compare to platform-reported ROAS. Update budget allocation accordingly.
Best practices for geo holdouts
- Match cells on pre-period sales, not population — you care about sales lift, not headcount.
- Avoid contaminated cells — DMAs adjacent to test cells with high cross-DMA shopping (e.g., NYC vs northern NJ) bleed exposure.
- Power-calculate the test in advance. If MDE (minimum detectable effect) is 30% and you expect 10% lift, you'll see noise, not signal.
- Run several cycles. One test = one data point. Confidence comes from repeated tests showing consistent lift.
- Document the test before launch. Pre-registration prevents post-hoc rationalization of inconclusive results.
User-level holdouts: ghost ads and PSA control
The cleanest methodology when the platform supports it. Amazon Brand Lift, Walmart Conversion Lift, Meta's Conversion Lift, Google's Conversion Lift all work the same way:
- Audience is built (e.g., in-market for category).
- Platform randomly splits eligible users into treatment and control.
- Treatment sees the real ad; control sees no ad (or, in some designs, a PSA).
- Platform tracks conversions for both groups.
- Lift = treatment conversion rate / control conversion rate — 1.
When to use
- When the platform supports it and your spend qualifies (often $25k–$100k+ minimum).
- When you need clean user-level causal inference for budget justification.
- When you want to measure new-to-brand vs repeat-buyer lift separately.
- When geo holdouts won't work (e.g., national e-commerce with no geographic targeting variation).
Media mix modeling for retail media
MMM has been around since the 1960s but the modern open-source revival (Meta's Robyn, Google's LightweightMMM, Uber's Orbit) plus dedicated vendors (Recast, Haus, Mass2 Analytics, Marketing Evolution) makes it accessible for mid-market brands. Apply MMM to retail media when:
- You have 2+ years of weekly sales data by retailer, ad spend by channel, plus seasonality, pricing, distribution, and macro indicators.
- You need cross-channel attribution that respects synergies (sponsored search depends on out-of-home awareness; off-site display lifts on-site conversions).
- You want to model adstock (carryover) and saturation (diminishing returns) properly.
- You need defensible numbers for board-level budget decisions.
MMM is not a replacement for experimentation. Best programs use MMM for strategic allocation and incrementality tests for tactical calibration of the MMM's assumptions.
Test design fundamentals
Power calculation
Before launching a test, calculate the minimum detectable effect (MDE) given your sample size, baseline conversion rate, and test duration. Standard inputs:
- Significance level: 95% (alpha = 0.05).
- Statistical power: 80% (beta = 0.20).
- Two-tailed test (you care about both lift and possible drag).
If your category has $50M annual sales in test geos and you can detect a 5% lift, that's a $2.5M effect — meaningful. If MDE is 25%, you'll only detect catastrophic effects, which isn't useful.
Duration
Calculate from MDE and traffic. For most retail media tests: 6–12 weeks for geo holdouts, 4–8 weeks for user-level holdouts on high-traffic audiences, 8–16 weeks for low-purchase-frequency categories.
Budget
A common design pitfall: holding back too much spend. If control geos go to zero and test geos go normal, you measure the all-or-nothing effect — useful but not actionable. Tests that vary spend levels (e.g., 50% spend vs 100% vs 150%) give you a response curve, which is much more useful for budget allocation.
Interpreting results
| Reported ROAS | iROAS | Interpretation |
| 10× | 8× | Healthy. Most reported value is real lift. |
| 10× | 3× | 30% real lift. Cannibalization is meaningful. Consider lower bid or pause on cannibalizing keywords/SKUs. |
| 10× | 0.5× | Almost entirely cannibalization. Pause or sharply reduce. |
| 5× | 6× | Reported understates impact. Halo effects on un-tracked SKUs. Increase budget. |
| 3× | Negative | Ad is suppressing baseline sales (rare but happens, e.g., when ad shows OOS product or competitor-loyal audiences). |
Advanced playbook
- Branded sponsored search incrementality test annually. The biggest budget hog with the worst iROAS-to-ROAS ratio. Annual test answers "how much would we lose if we paused?" The answer is usually less than you think.
- Run incrementality on new audiences before scaling. Before pouring budget into a new lookalike audience or competitor conquest, run a 4-week ghost-ad test to confirm causal lift.
- Combine geo and ghost-ad tests for triangulation. Independent methodologies converging on the same iROAS estimate is much stronger evidence than either alone.
- Test promotional moments separately. Black Friday iROAS is structurally different from Tuesday-in-March iROAS. Don't average them.
- Halo measurement. Don't just measure ad-tagged SKU sales. Measure brand-level sales lift. Halo is real and often equals or exceeds direct lift.
- New-to-brand-specific tests. Optimize for NTB% separately. The audience that yields highest NTB% may not yield highest ROAS, but the long-term LTV math may justify it.
- Cross-retailer cannibalization tests. If you advertise on Amazon and Walmart, run a Walmart-pause test in select geos to see if Amazon sales rise to compensate (cannibalization) or stay flat (genuine incremental).
- MMM-experiment hybrid. Use experiment results to calibrate MMM priors (Bayesian MMM frameworks like Meta's Robyn support this directly).
- Pre-registration discipline. Document hypothesis, MDE, duration, success criteria before launch. Forces honesty.
- Build an incrementality cadence. Quarterly geo holdout on largest retailer, annual user-level holdout on big audiences, MMM rebuilt twice a year, ad-hoc tests for major budget decisions.
Common mistakes
- Trusting platform-reported ROAS as the only number that matters.
- Skipping power calculation; running underpowered tests that fail to find effects that exist.
- Pausing branded sponsored search across all geos "to see what happens" without test design.
- Running tests on noisy weeks (holiday spikes, supply disruptions, news events) without controls.
- Cherry-picking favorable cells post-hoc; rationalizing inconclusive results.
- Running one test once and treating the number as truth.
- Confusing observational lift studies with causal incrementality — they aren't the same.
- Letting the retailer's own "lift study" replace independent measurement — the retailer is not a neutral party.
- Ignoring cross-retailer cannibalization.
- Optimizing for iROAS alone and missing brand-building / new-to-brand value.
Operating checklist
- Quarterly geo holdout test cadence on top retailer
- Annual branded sponsored search incrementality test
- Power calculations documented before every test launch
- Pre-registration of hypothesis, MDE, duration, success criteria
- User-level holdouts (Conversion Lift / Brand Lift) on largest audiences quarterly
- MMM rebuilt twice yearly; calibrated against experimental results
- Reported ROAS and iROAS shown side-by-side in monthly reporting
- NTB% tracked alongside iROAS
- Cross-retailer cannibalization test annually
- Documented decision rule: budget shift threshold based on iROAS change
Sources and further reading
- Amazon Brand Lift & Amazon Marketing Cloud measurement documentation
- Walmart Connect Conversion Lift methodology
- Meta Conversion Lift and Google Conversion Lift methodology
- Catalina Sales Lift studies — CPG retail media incrementality benchmarks
- Circana (formerly IRI) and NielsenIQ MMM methodology references
- Meta Robyn, Google LightweightMMM, Uber Orbit — open-source MMM frameworks
- Recast, Haus, Mass2 Analytics, Marketing Evolution — commercial MMM and incrementality vendors
- Andrew Stephen et al., "Effective Retail Media Measurement" — Saïd Business School research
- Avinash Kaushik, Occam's Razor — experimentation methodology
- Andrew Gelman, Bayesian Data Analysis — statistical foundations
- Wes Nichols, "Path to Purchase" — cross-channel attribution methodology
- The Drum, AdExchanger, MarTech — coverage of measurement evolution
Part of the Retail Media series. Continue to the next module or take the series exam.