Measuring incrementality in retail ads: geo tests and holdouts on a small budget

Ask any ad platform how its campaigns performed and it will tell you. Ask whether those campaigns changed anything, and the answer gets quiet. That gap between reported return on ad spend and actual added revenue is where most retail media budgets quietly leak, and it is the gap that incrementality testing exists to close.

The good news for a small retailer is that the method does not require a data science team or a seven-figure budget. A geo split or a clean holdout can be run with a spreadsheet, a calendar, and the discipline to leave one group of customers alone for a few weeks. The hard part is not the math. It is resisting the urge to switch the ads back on when the test markets look soft in week two.

In short

  • Attributed ROAS is not incremental ROAS. Platform reporting credits conversions that would have happened without the ad, especially in brand search and retargeting.
  • A holdout is the cheapest honest measurement you can buy. You withhold ads from a randomly chosen group and compare outcomes against the group that still sees them.
  • Geo splits solve the user-level problem. When platforms will not let you randomize by person, you randomize by city, region or postcode instead.
  • Small budgets can only detect large effects. A shop doing $40,000 a week in revenue cannot reliably measure a 3% lift, so design the test around what you can actually see.
  • The decision rule matters more than the p-value. Write down before the test what result would make you cut, hold or scale the channel, then follow it.

Why attributed ROAS overstates paid performance

Every major ad platform runs its own attribution model, and every one of them is both the referee and a player. When a shopper clicks a retargeting ad for a product already in their cart, the platform records a conversion. It cannot record the counterfactual, which is that most of those shoppers were going to return anyway. The result is systematic overcounting, and the size of the overcount varies enormously by channel.

This is not a fringe concern. The practice of separating correlation from causation in advertising measurement sits at the heart of the broader shift away from last-click reporting, a shift covered in more depth in our guide to retail marketing in the age of AI search and social commerce. Attribution answers “which touchpoint was nearest the sale”. Incrementality answers “would the sale have happened at all”.

The brand search trap

Brand search is the classic example. A customer types your shop name into Google, sees your paid ad above your own organic listing, clicks the ad, and buys. The platform books a conversion at a spectacular ROAS. In reality the customer had already decided to visit you, and the organic result sitting directly below would have caught the same click for free.

Controlled brand search experiments have repeatedly found that a large share of paid brand clicks cannibalize organic clicks rather than adding new ones. The exact share depends on how strong your organic position is, how many competitors bid on your name, and how crowded the results page has become. A retailer sitting at position one organically with no competitor bidding has the least to gain. A retailer whose name is being bid on by three marketplaces has the most.

Retargeting counts loyal buyers twice

Retargeting pools are built from people who already visited your site, browsed products and often abandoned a cart. That is, by construction, your highest-intent audience. Showing them an ad and claiming credit for their purchase is close to claiming credit for the sunrise. Retargeting can genuinely work, particularly for considered purchases with long decision windows, but the measured ROAS in platform reporting is almost never the incremental one.

The practical effect for a small retailer is a budget that drifts toward the channels that report best rather than the channels that perform best. Prospecting campaigns, which are the only ones that can actually bring new customers, look weak on platform dashboards and get cut first. The specifics of where that bias bites hardest across channels are worked through in our overview of paid ads for retailers in 2026.

Why the platforms are not lying, exactly

It helps to be precise about the failure. Platforms report what their models observe, and those models are usually internally consistent. The problem is that an attribution model is a bookkeeping convention, not an experiment. It assigns credit among touchpoints that occurred; it never constructs the world in which the ad did not run.

Signal loss has made this worse. Since the rollout of app tracking transparency on iOS and the long tail of browser cookie restrictions, platforms have increasingly filled measurement gaps with modeled conversions. Modeled numbers are estimates trained on the observable subset, and they inherit the same inability to see the counterfactual. Our playbook for Meta retail ads after the iOS privacy shift covers how that modeling changed what the dashboards mean.

Holdout tests explained in retail terms

A holdout test borrows the logic of a randomized controlled trial. You split your addressable audience into two groups at random. One group, the treatment group, sees your ads as normal. The other group, the holdout or control group, is deliberately excluded. After a set period you compare revenue per person between the groups. The difference is the incremental effect.

The appeal is that randomization handles every confounder you have not thought of. You do not need to model seasonality, competitor activity, weather or a viral post, because whatever happened, happened to both groups. The cost is real but bounded: you give up the sales the holdout would have generated from ads, for the duration of the test.

What a holdout actually measures

Suppose you run a 4-week holdout across 100,000 customers, split evenly. The treatment half generates $180,000 in revenue. The holdout half generates $150,000. The incremental revenue attributable to advertising is $30,000. If you spent $20,000 reaching the treatment half, your incremental ROAS is 1.5. Your platform may well have reported 4.2 for the same spend.

Notice what that number lets you do. A reported ROAS of 4.2 tells you nothing actionable, because you have no idea what the baseline was. An incremental ROAS of 1.5 compares directly against your gross margin. If your blended margin is 45%, then $1.50 of incremental revenue per dollar spent returns about 68 cents of gross profit per dollar, and the channel is losing money at that spend level.

Reading iROAS instead of ROAS

The breakeven calculation is the whole point of the exercise. Incremental ROAS multiplied by contribution margin needs to exceed 1.0 for the channel to pay for itself in-period. Retailers with strong repeat purchase rates can justify a lower in-period threshold, because the first order carries lifetime value behind it, but that argument should be backed by actual cohort retention data rather than optimism.

Metric What it measures Who computes it Typical bias
Platform ROAS Conversions credited to the ad by the platform model The ad platform Overstates, often heavily on brand and retargeting
Blended ROAS Total revenue divided by total ad spend You Mixes organic baseline into paid performance
Incremental ROAS Added revenue caused by the ad, versus a control An experiment Unbiased if randomized, noisy if underpowered
Marketing mix model output Channel contribution inferred from historical variation A statistical model Depends on model assumptions, needs long history
Last-click conversions Sales where the ad was the final touch Analytics tool Ignores upper funnel entirely

Geo splits when you cannot split by user

User-level holdouts are clean, but they are not always available. Some platforms only offer them at high spend tiers. Others cannot suppress ads for a specific list reliably, particularly in broad reach campaigns and in automated formats where targeting control is deliberately limited. In those cases you randomize geography instead.

A geo split assigns markets rather than people. You might pick 20 comparable metro areas, randomly assign 10 to keep running ads and 10 to go dark, then compare total revenue by market over the test window. The unit of randomization becomes the city, which has consequences for statistical power that catch people out.

Choosing matched markets

Random assignment works best when the pool is large. With only 20 markets, pure randomness can easily hand you all the big ones on one side. Matched-market design fixes this by pairing markets on pre-period behavior first, then randomizing within each pair. Pair Columbus with Kansas City, Tucson with Fresno, and so on, matching on historical revenue, order volume, growth trend and seasonality shape rather than on population alone.

Pre-period correlation is the test of a good pair. Plot weekly revenue for both markets over the previous 6–12 months. If the two lines track each other closely, the pair will hold during the test. If one spikes every December and the other does not, you have a seasonality mismatch that will show up as a fake effect.

Designated market areas and the online problem

Geo testing was invented for broadcast media, where designated market areas map neatly onto transmission footprints. Online retail is messier. Customers order from one postcode and ship to another, mobile traffic carries imprecise location, and a customer who sees an ad while travelling does not stay in the test cell. These leakages all push measured effects toward zero, which means a geo test is conservative: it is more likely to miss a real effect than to invent one.

For retailers with physical stores the mapping is cleaner, because catchment areas are real and purchases are geographically anchored. For pure online sellers, consider using shipping region rather than IP-derived location as the assignment key, and exclude markets that sit on the boundary between two test cells.

Which design fits which situation

Design Randomization unit Best for Minimum practical spend Main weakness
User-level holdout Individual customer Retargeting, email-matched audiences, CRM campaigns Low, if the platform supports it Not offered on every channel or format
Matched-market geo split City or region Broad prospecting, local retail, store traffic Medium Needs many markets for power
Public service announcement (PSA) test Individual customer Display and video where suppression is hard High, you pay for the control ads You spend money on ads that cannot convert
Switchback or on/off scheduling Time period Single-market businesses, delivery and marketplace apps Low Carryover between periods contaminates the result
Marketing mix model None, observational Budget allocation across many channels at once None, but needs 2+ years of data Correlational, sensitive to specification

Sizing a test your budget can actually power

This is the section most guides skip, and it is the one that decides whether your test produces an answer or a shrug. Statistical power is the probability that your test detects a real effect when one exists. Underpowered tests return inconclusive results, and inconclusive results get interpreted as “the ads do not work”, which is a different and often wrong conclusion.

Power depends on three things: the size of the effect you are trying to detect, how noisy your revenue is week to week, and how many units (people or markets) you have. You control the third only partly. You can control the first by deciding in advance what size of effect would actually change your behavior.

Start from the decision, not the data

Ask what lift would make you keep spending. If your paid channel costs 8% of revenue and your contribution margin is 40%, the channel needs to add roughly 20% to the revenue of exposed customers to break even. That is your minimum detectable effect. There is no point designing a test that can only see a 2% lift when the decision hinges on whether the lift is above or below 20%.

This reframing is liberating for small budgets. A shop that needs to distinguish “huge effect” from “almost nothing” needs far less data than a brand trying to measure a 3% difference, and most small retailers are in the former situation whether they realize it or not.

A rough sizing table

The figures below are illustrative rules of thumb for a two-sided test at conventional 80% power, assuming weekly revenue variation typical of a specialty retailer. Treat them as a starting point for conversation with whoever runs your analytics, not as a substitute for a power calculation on your own historical data.

Weekly revenue in test Weeks needed for 20% lift Weeks needed for 10% lift Weeks needed for 5% lift Verdict
$10,000 6–8 Impractical Impractical Test only large effects
$50,000 3–4 10–14 Impractical Quarterly cadence is realistic
$250,000 2 4–6 16+ Can test channel by channel
$1,000,000 1–2 2–3 6–8 Continuous testing feasible

The pattern is the important part. Halving the effect you want to detect roughly quadruples the data you need. A retailer who insists on measuring small effects on a small budget will run tests forever and never conclude anything.

Test duration, seasonality and noise

Four weeks is the common default, and it is a reasonable default for a reason. It covers a full monthly pay cycle, smooths day-of-week variation, and gives slower purchase decisions time to land. Shorter than two weeks and you are mostly measuring noise. Longer than eight and the business changes underneath you.

Purchase cycle length should drive the decision more than convention does. If your average time from first ad exposure to order is 3 days, a 2-week test captures nearly all the effect. If you sell furniture with a 6-week consideration window, a 4-week test will see only a fraction of the eventual impact and will understate the channel badly.

What to avoid scheduling around

Do not run a test that straddles a major promotional period unless the promotion is the thing you are testing. Black Friday, a clearance event or a product launch will dominate the variance and swamp the ad effect. Equally, avoid windows where you plan to change pricing, email cadence, or site experience, because those changes will not respect your test cells.

Holidays create a subtler problem in geo tests. Regional holidays, school calendars and local events hit markets asymmetrically. A test running through a week when half your treatment markets had a major local festival will produce a result that has nothing to do with advertising. Check the calendar for every market in the design before you start.

Carryover and the cooldown period

Advertising effects do not stop the moment the campaign does. Customers who saw an ad in week one may buy in week five. In a geo test this means the dark markets are not truly clean on day one, especially if you have been advertising there heavily for months. A one to two week cooldown before measurement begins reduces the contamination, at the cost of a longer overall test.

The same logic applies after the test ends. Turning ads back on in the holdout markets produces a catch-up spike as deferred demand converts. That spike is not incremental growth; it is timing. Exclude the first week after resumption from any before-and-after comparison you make.

Reading the result without fooling yourself

You now have two numbers, revenue in treatment and revenue in control, and a strong temptation to declare victory. Resist it for a moment, because the most common failure in small-budget testing is not a bad experiment. It is a good experiment read badly.

Validate the pre-period first

Before looking at the test window, run the same comparison on the period before the test started, when both groups were treated identically. The measured “effect” in that window should be approximately zero. If the pre-period shows a 9% difference between your cells, your randomization did not balance, and the test window result is not interpretable. This single check catches more broken geo tests than anything else.

Report an interval, not a point

An incremental ROAS of 1.8 sounds precise. An incremental ROAS of 1.8 with a 95% confidence interval spanning 0.4 to 3.2 tells you the honest truth, which is that you have learned the channel is probably not catastrophic and probably not transformative. Always compute and report the interval. Decisions made on point estimates from noisy tests are how budgets whipsaw.

Where the interval sits relative to your breakeven threshold is the real output. If the entire interval sits above breakeven, scale with confidence. If the entire interval sits below, cut. If it straddles breakeven, hold the current spend and either run longer or accept the ambiguity rather than pretending it resolved.

Do not slice until something turns significant

Once the headline result disappoints, the temptation is to look at the subset where it worked: new customers only, one product category, mobile traffic, the three markets that did well. Each slice you examine raises the chance of finding an apparent effect that is pure noise. Decide your primary metric and your segments before you see the data, write them down, and treat everything else as a hypothesis for the next test rather than a finding from this one.

Turning findings into budget decisions

A test that does not change a budget was a waste of four weeks. The handoff from result to decision needs to be written before the test runs, in plain language, with specific numbers. Something like: if incremental ROAS lands above 2.4 we increase this channel’s budget by 30% next quarter; if it lands below 1.2 we cut it in half and move the money to prospecting; between those we hold.

Apply the finding where it generalizes

One test measures one channel, at one spend level, in one season, for one product mix. Diminishing returns are real, so a channel that returns an incremental ROAS of 2.5 at $20,000 per month will not return 2.5 at $80,000. The honest use of a single test is directional: it tells you which way to move, not how far.

Repeat the test at the new spend level once you have moved. Over a year, three or four tests at different spend levels sketch out a response curve, and a response curve is what actually lets you set budget rather than guess at it. Retailers running automated campaign types face an extra wrinkle here, because the system reallocates spend internally, a dynamic covered in our breakdown of Performance Max for retailers.

Feed the result back into your attribution

The most valuable long-term output of incrementality testing is a calibration factor. If your test says the true effect is 40% of what the platform reported, you can apply that discount to the platform’s daily numbers between tests and run on something closer to reality. Recalibrate every two or three quarters, or whenever the platform changes its measurement methodology.

Calibration also helps when comparing channels that report in incompatible ways. A shopping campaign and a social prospecting campaign measured on their own dashboards are not comparable at all. Measured against a common incremental baseline, they finally are. The mechanics of the shopping side, including how feed quality drives what the dashboard shows, are covered in our primer on Google Shopping ads for retail beginners.

Build the cadence into the calendar

Incrementality testing works as a habit, not an event. A practical cadence for a mid-sized retailer is one test per quarter, rotating across channels, with brand search tested at least annually because the cannibalization rate moves as competitors enter and exit. Put the dates in the calendar at the start of the year, because a test that has to be justified from scratch each time never gets run.

This discipline is what separates retailers who know what their marketing does from those who have a dashboard. Context for how incrementality fits alongside organic search, owned channels and emerging AI-driven discovery is laid out in our retail marketing guide, which is worth reading alongside this piece. For broader context on how much of US retail trade now runs through online channels, the quarterly series published by the US Census Bureau is the standard reference.

FAQ on incrementality testing

What is the difference between attribution and incrementality?

Attribution divides credit for a sale among the touchpoints that preceded it. Incrementality asks whether the sale would have occurred without the advertising at all. A channel can receive high attribution credit and have close to zero incremental effect, which is exactly what happens with most brand search and much retargeting.

How much revenue do I give up by running a holdout?

The cost equals the incremental revenue the holdout group would have produced, which is by definition the thing you are measuring. If you hold out 20% of your audience for 4 weeks and the true incremental effect is 10%, you forgo roughly 2% of revenue for a month. Many retailers find that smaller than the waste the test uncovers.

Can I run an incrementality test with a $5,000 monthly ad budget?

Yes, but only for large effects. At that spend level you can reliably distinguish “this channel does something substantial” from “this channel does almost nothing”, which is often the decision you actually face. You cannot measure a 5% lift, and attempting to will produce an inconclusive result you may misread as a negative one.

How many geographic markets do I need for a geo split?

Twenty markets is a common practical floor, and more is substantially better because the market is the unit of randomization. Below roughly ten markets, a single unusual city can drive the entire result. If you operate in fewer markets than that, a switchback design across time periods is usually the more honest option.

Which channel should I test first?

Test wherever the gap between reported performance and plausible reality is widest, which for most retailers means branded search or retargeting. Those channels usually carry the highest reported ROAS and the lowest true incrementality, so a test there has the largest chance of freeing up budget.

What if the test shows no significant effect?

Check the power of the design before concluding the channel is worthless. A non-significant result in an underpowered test means you learned nothing, not that the effect is zero. Look at the confidence interval: if it spans everything from heavy loss to strong profit, the test was too small, and the correct response is a longer or larger follow-up rather than a budget cut.

Do platform-run lift studies count as real incrementality tests?

They use genuine experimental designs and are considerably better than attribution reporting, so they are worth running when available. The caveat is that the platform designs, executes and reports the study on its own inventory, so an independently run geo test provides a useful cross-check on the result. Treat the two as complementary rather than interchangeable.

How often should a retailer repeat these tests?

Quarterly is a reasonable cadence for an active advertiser, rotating the channel under test each time. Repeat sooner if you change spend materially, launch a new product line, or the platform announces a measurement or targeting change, since any of those invalidates the calibration you derived from the previous test.

Can incrementality testing replace a marketing mix model?

They answer different questions and work best together. Experiments give unbiased answers about one channel at one spend level; mix models allocate across the whole portfolio but rest on assumptions. The strongest setup uses experimental results to calibrate and validate the mix model, so the model’s channel coefficients are anchored to something causally measured.