HomeTrainingRetail Media › Retail Media Incrementality
Retail Media
RGM° · Training

Retail Media Incrementality

Why reported ROAS lies. The methodologies, test designs, and decision rules that separate real lift from cannibalization.

What you will learn

  1. Why retail media ROAS lies and incrementality is the truth
  2. The three types of incrementality testing
  3. Geo holdouts: the workhorse methodology
  4. User-level holdouts: ghost ads and PSA control
  5. Media mix modeling for retail media
  6. Test design: power, duration, MDE, budget required
  7. Interpreting results: what counts as a win
  8. Advanced playbook
  9. Common mistakes
  10. Operating checklist

Why retail media ROAS lies

Sponsored search on Amazon, Walmart, or Target reports 8×, 12×, or 20× ROAS routinely. Brand teams celebrate, scale spend, and the platforms report even better ROAS. Then quarterly sales review shows total category growth is flat or even declining despite ad spend tripling. What happened?

The answer is structural. Platform-reported ROAS is last-click attribution to shoppers already at the retailer searching for relevant products. A large share of those shoppers would have purchased anyway — especially on branded searches. The ad gets credit for purchases it didn't cause.

Industry research (Amazon's own MMM studies, Catalina case studies, Nielsen iROAS benchmarks) suggests that incremental ROAS (iROAS) is often 30–70% of reported ROAS. For branded sponsored search, iROAS can be as low as 10–20% of reported. For category prospecting and new-to-brand, iROAS can exceed reported because of halo effects on related SKUs.

This is not unique to retail media — the same dynamic applies to branded paid search on Google — but the scale of cannibalization is larger in retail media because the user's purchase intent is so high at the moment of ad exposure.

The three types of incrementality testing

TypeHow it worksStrengthsLimits
Geo holdoutPause or reduce ads in selected geos; compare sales in test vs control geosWorks without user-level data; controllable; respects privacyRequires sufficient geo separation; can't isolate audience-level effects
User-level holdout (ghost ads / PSA)Random users in matched audience see ad; matched control sees public-service ad or no ad; compare conversionsClean causal inference at user level; tight statistical controlRequires platform support; small audience sizes need long test duration
Media mix model (MMM)Statistical model fitted to 2+ years of weekly sales, spend, and external factors; isolates channel contributionCross-channel attribution; accounts for adstock, saturation, haloSlow (rebuild quarterly), expensive, requires data discipline

Geo holdouts: the workhorse methodology

Geo holdouts are the most common incrementality methodology in retail media. The mechanics:

  1. Select test cells. Group DMAs or zip codes into 2–4 matched cells based on historical sales, geography, and demographics. Tools like Google's Geo Experiments framework or third-party tools (Haus, Recast, Lift) help.
  2. Assign treatment. Cells get treatment (ad on / off / spend levels). Common designs: A/B (50/50 split), A/A/B (33/33/33), or matched-pair (one cell heavy, matched cell light, third cell control).
  3. Run for sufficient duration. 6–12 weeks typical. Daily-purchase categories can be shorter; quarterly-purchase categories need longer.
  4. Measure lift. Compare sales (from retailer point-of-sale or syndicated data) across cells. Apply statistical test (t-test, regression, or causal-inference framework).
  5. Translate to iROAS. Lift / spend = iROAS. Compare to platform-reported ROAS. Update budget allocation accordingly.

Best practices for geo holdouts

User-level holdouts: ghost ads and PSA control

The cleanest methodology when the platform supports it. Amazon Brand Lift, Walmart Conversion Lift, Meta's Conversion Lift, Google's Conversion Lift all work the same way:

  1. Audience is built (e.g., in-market for category).
  2. Platform randomly splits eligible users into treatment and control.
  3. Treatment sees the real ad; control sees no ad (or, in some designs, a PSA).
  4. Platform tracks conversions for both groups.
  5. Lift = treatment conversion rate / control conversion rate — 1.

When to use

Media mix modeling for retail media

MMM has been around since the 1960s but the modern open-source revival (Meta's Robyn, Google's LightweightMMM, Uber's Orbit) plus dedicated vendors (Recast, Haus, Mass2 Analytics, Marketing Evolution) makes it accessible for mid-market brands. Apply MMM to retail media when:

MMM is not a replacement for experimentation. Best programs use MMM for strategic allocation and incrementality tests for tactical calibration of the MMM's assumptions.

Test design fundamentals

Power calculation

Before launching a test, calculate the minimum detectable effect (MDE) given your sample size, baseline conversion rate, and test duration. Standard inputs:

If your category has $50M annual sales in test geos and you can detect a 5% lift, that's a $2.5M effect — meaningful. If MDE is 25%, you'll only detect catastrophic effects, which isn't useful.

Duration

Calculate from MDE and traffic. For most retail media tests: 6–12 weeks for geo holdouts, 4–8 weeks for user-level holdouts on high-traffic audiences, 8–16 weeks for low-purchase-frequency categories.

Budget

A common design pitfall: holding back too much spend. If control geos go to zero and test geos go normal, you measure the all-or-nothing effect — useful but not actionable. Tests that vary spend levels (e.g., 50% spend vs 100% vs 150%) give you a response curve, which is much more useful for budget allocation.

Interpreting results

Reported ROASiROASInterpretation
10×Healthy. Most reported value is real lift.
10×30% real lift. Cannibalization is meaningful. Consider lower bid or pause on cannibalizing keywords/SKUs.
10×0.5×Almost entirely cannibalization. Pause or sharply reduce.
Reported understates impact. Halo effects on un-tracked SKUs. Increase budget.
NegativeAd is suppressing baseline sales (rare but happens, e.g., when ad shows OOS product or competitor-loyal audiences).

Advanced playbook

Common mistakes

Operating checklist

Sources and further reading


Part of the Retail Media series. Continue to the next module or take the series exam.