Measurement September 28, 2026 15 min read

Incrementality On Amazon: Proving Your Ads Caused The Sale

Every article tells you to pause branded search and see what happens. On Amazon that test degrades your organic ranking while it runs, which corrupts the measurement and costs you money after it ends.

1 Question ROAS Cannot Answer
4 Test Designs Available
2-4 Weeks A Test Typically Runs
0 Tests Worth Running Without A Plan
Quick Answer

Attribution is bookkeeping and incrementality is causation. A campaign's ROAS tells you which sales were credited to it, not which sales would have failed to happen without it, and those are very different numbers. On Amazon the honest way to close that gap is an experiment: hold a comparable group back from your ads and measure the difference. Four designs are practical, being an audience holdout inside DSP, a geographic holdout, a time-based pause on a specific campaign or ad type, and platform-run lift studies. The complication that most guidance ignores is Amazon-specific. Amazon's ranking responds to sales velocity, so pausing ads suppresses your organic position while the test runs. That contaminates the result, because you are measuring ads-off plus a decaying organic baseline rather than ads-off alone, and it leaves a real cost behind after the test ends. Design around it by testing on a subset of products rather than your whole catalog, keeping the window short, setting your confidence threshold before launch, and accepting that some questions are not worth the ranking damage to answer.

A shopper types your brand name, sees your ad, clicks it, and buys. Your report records a conversion. It cannot record the fact that they would have scrolled two inches down and bought anyway.

That single scenario is the whole problem. Branded search campaigns routinely report eight or ten times return on ad spend, which makes them look like the best-performing thing in the account. They are also the campaigns most likely to be buying sales that were already yours, because the shopper who searched your brand name had already decided who they were shopping for.

Nothing in your advertising console can distinguish those two cases. Attribution answers which touchpoint got credit. It has no mechanism for answering what would have happened otherwise, because the counterfactual is not in the data.

The only way to get at it is to run an experiment. This post covers how to do that on Amazon specifically, including the complication that makes Amazon harder than Meta or Google.

01/12 Section

Why ROAS Cannot Answer This

The distinction is worth stating precisely because it is frequently blurred by people who understand it perfectly well.

Attribution assigns credit for a conversion that occurred, according to rules about which touchpoints count and over what window. It is an accounting exercise performed on things that happened. Incrementality asks a different question: how many of those conversions would have occurred without the advertising?

Definition

Incrementality. The causal contribution of advertising, measured as the difference in outcomes between a group exposed to ads and a comparable group deliberately withheld from them. It answers what the advertising added rather than what it was credited with. Incremental return on ad spend divides that measured lift by the spend producing it, which is why it is typically lower, sometimes far lower, than platform-reported return.

The reason this matters commercially is that budget decisions made on attributed ROAS are systematically distorted in a predictable direction. Campaigns capturing existing demand look excellent. Campaigns creating new demand look mediocre. Money flows toward the former, which is exactly backwards if you are trying to grow.

None of this makes ROAS useless. It is a fine operational metric for comparing two similar campaigns or spotting a break. It is a poor instrument for deciding whether a channel deserves to exist, and it is used for that constantly.

02/12 Section

Where Attribution Overstates Most

Overstatement is not evenly distributed. It concentrates in four places, and knowing which they are tells you where testing is worth the effort.

Campaign TypeWhy Attributed ROAS Flatters ItTest Priority
Exact-match branded searchThe shopper already chose you. Organic listing would likely have captured the sale.Highest. Usually the largest gap.
RetargetingTargets people already far down the funnel, many of whom were returning regardless.High.
Loyal-customer audiencesReplenishment buyers on consumables repurchase on their own schedule with or without an ad.High for subscription-type products.
Overlapping campaignsSeveral campaigns targeting the same shopper each claim the same conversion.Medium. Fix exclusions first.
Category and competitor termsFrequently understated rather than overstated. Reaches shoppers who were not looking for you.Low, but worth confirming.

The last row is the one to sit with. If branded search is overstated and category acquisition is understated, then a brand optimizing to attributed ROAS is steadily moving budget away from growth and toward credit-taking, while its dashboard improves the whole time.

That pattern is worth checking before you test anything, because campaign overlap in particular has a cheaper fix. If multiple campaigns are serving the same shoppers without exclusion logic, tightening the structure resolves double-counting without an experiment. Our Amazon PPC strategy guide covers that structural work.

03/12 Section

What Incrementality Actually Means

One source puts it well: most brands that claim to test incrementality are actually measuring attribution shift. The distinction is worth guarding.

Turning off a campaign and watching attributed sales move between campaigns is not an incrementality test. That is credit relocating. A real test requires two comparable groups, one exposed and one deliberately withheld, and measures the difference in total outcomes between them, not the difference in what got credited.

The comparison that matters is total units or total revenue across the whole business for those groups. If pausing a campaign leaves total sales unchanged while attributed sales collapse, you have learned something valuable: the campaign was relabelling demand rather than creating it.

Not The Same As Listing Tests

This is distinct from experimenting on images, titles, or A+ content. Those tests ask which version converts better among people who arrived. Incrementality asks whether the media that brought them was necessary at all. Both are worth doing and they answer different questions. Our guides to Amazon's experimentation tools and ecommerce A/B testing cover the listing side.

04/12 Section

The Four Test Designs

Each has different requirements and different weaknesses. Pick on what you can actually run rather than what is theoretically cleanest.

Designs Available On Amazon Pick For Feasibility
Design 01
Audience Holdout

A slice of the addressable audience is withheld from ads inside DSP. The cleanest randomisation available, and it requires a DSP entity. AMC reporting supports the holdout logic.

Design 02
Geographic Holdout

Suppress a channel in matched regions and compare against the rest. Works without DSP, but needs regional sales data and enough volume per region to read.

Design 03
Product Subset Pause

Pause a campaign type on a controlled subset of products and compare against a matched set. Practical for most sellers and where the ranking risk concentrates.

Design 04
Platform Lift Study

Platform-run randomisation into exposed and control. Convenient, and the platform grades its own homework, so read the methodology before trusting the output.

A note on the fourth. Platform lift studies have a known measurement asymmetry in the wider industry, where exposed groups are tracked more completely than control groups, producing apparent lift that partly reflects better observation rather than better outcomes. That criticism is documented mainly against other platforms, and the general principle stands: a measurement run by the party being measured deserves scrutiny.

For most Amazon sellers without DSP, design three is the realistic option, and it is also the one where the Amazon-specific complication in section six bites hardest.

05/12 Section

Sizing The Test

Thin spend produces thin signal. If your test cannot move enough volume to register a measurable gap between exposed and control, you cannot separate effect from noise, and you will have paid for an inconclusive result.

Decide the following before launch rather than after, because deciding afterward is how brands talk themselves into a conclusion the data does not support.

  1. The confidence threshold. Commonly 95 percent. Set it in writing before you start.
  2. The minimum effect worth detecting. If a 3 percent lift would not change your decision, do not design a test capable only of finding 3 percent.
  3. The duration. Two to four weeks is the typical window. Shorter rarely accumulates enough volume; longer accumulates too much contamination.
  4. The primary metric. Total units or total revenue across the business, not attributed sales. Write this down or someone will report the wrong one.
  5. What you will do with each outcome. Both directions. A test with only one actionable result is not a test.

Geographic tests carry additional data requirements. One source describing geo methodology calls for first-party performance data at least at state level, reported weekly or daily, with at least six months of clean historical data before the test to establish baselines. That is a meaningful prerequisite and many brands do not have it.

06/12 Section

The Ranking Trap

Here is what makes Amazon different from every other platform you might run this on, and it is largely absent from generic incrementality guidance.

Amazon's organic ranking responds to sales velocity. Products that sell well rank better, and ranking better produces more sales. Advertising feeds that loop, because ad-driven sales are still sales.

So when you pause advertising on a product to measure incrementality, two things happen simultaneously. Ad-driven sales stop, which is what you intended to measure. And the product's sales velocity drops, which degrades its organic position, which reduces organic sales, which degrades velocity further.

On Meta, pausing ads removes a source of traffic. On Amazon, pausing ads removes a source of traffic and simultaneously weakens the organic position that was supplying the rest of it.
Why a branded search pause test overstates incrementality on Amazon

The measurement consequence is specific and it runs in one direction. Your test will show a larger sales drop than the advertising alone caused, because part of the decline came from ranking decay triggered by the pause. You will conclude the ads were more incremental than they were.

The commercial consequence is worse and outlasts the test. Ranking lost during a pause does not return the moment you resume. You have to rebuild velocity to recover position, which means the true cost of the experiment includes a recovery period nobody budgets for. One source refers to needing patience to weather the short-term ranking dip, which is a mild way of describing a real expense.

For how the ranking mechanism works in detail, our breakdown of Cosmo against the older A9 and A10 models covers the velocity relationship this exploits.

07/12 Section

Designing Around It

The ranking effect cannot be eliminated, and it can be reduced enough that a test is worth running.

Test on a subset, never the whole catalog. Pause on a controlled group of products and compare against a matched set that keeps running. This contains the ranking damage to products you chose to risk, and the matched control also absorbs seasonality and category movement.

Choose products with established rank. A product with a long stable ranking history has more resilience than a recent launch still building position. Never run this on a product in its ramp period.

Keep the window short. Two weeks limits decay compared with four. There is a genuine trade against statistical power, and on Amazon the shorter end is usually the better compromise.

Prefer audience holdouts where you can. A DSP audience holdout withholds ads from a slice of people rather than removing ads from a product entirely, so aggregate velocity holds up and ranking is largely protected. This is the strongest argument for running incrementality work through DSP if you have it.

Reduce rather than eliminate. Cutting spend substantially instead of going dark preserves some velocity while still creating a measurable gap. Less clean statistically, considerably cheaper in ranking terms, and often the right trade.

Measure the recovery too. Track how long ranking takes to return after the test. That number is part of the cost of testing and you will want it before deciding whether to test again.

08/12 Section

Reading The Result Honestly

The failure mode after a test is not usually bad arithmetic. It is over-claiming.

Incremental ROAS is the measured revenue lift divided by the spend that produced it. If pausing spend cost you materially less revenue than the platform claimed that spend generated, your incremental figure is lower than the reported one, and the correct comparison for a budget decision is incremental ROAS against your contribution-margin breakeven rather than against the reported number.

Three honesty checks before you present anything.

  • State the confidence interval, not the point estimate. A result of 2.1x with a range spanning 0.8x to 3.4x is a different finding from 2.1x with a tight band, and only one of them justifies a decision.
  • Say what was confounded. On Amazon that means naming the ranking effect explicitly and acknowledging your measured lift is inflated by it.
  • Do not generalise beyond the test. A result on ten products in one category during three weeks in September is a result on ten products in one category during three weeks in September.

The single most useful discipline is to write down what you expect before you look. A prediction recorded in advance turns a result into evidence. A prediction formed after seeing the number is a story.

Free Resource

The Ecom Profit Box

Our collection of ecommerce growth resources, including the margin frameworks a test result has to be judged against.

Get It Free
30 Minutes

Is Your Branded Spend Working?

We will look at your structure and tell you whether a test is warranted or whether cheaper fixes come first.

Book A Call
09/12 Section

The Benchmark Numbers

Several figures circulate in this space. They are directionally useful and individually unciteable, and the distinction matters.

Circulating FigureSource TypeHow To Use It
Branded search: 1.5x to 3x incremental against 10x+ attributedMeasurement vendorAs a hypothesis to test on your own account, not as your number.
Retargeting: 1.5x to 2.5x incrementalMeasurement vendorSame. Direction is credible, magnitude is not yours.
iROAS 40% to 70% below platform-reported across DTCVendor blog, described as research, no citation givenTreat as folklore with a plausible mechanism behind it.
A grocery chain found 0% lift from non-branded paid search in 12 marketsSingle vendor case studyIllustrates that zero is a possible answer. Says nothing about your category.

The last one is worth keeping for one reason. It demonstrates that a properly run test can return no measurable lift at all, and that this is a finding rather than a failure. The chain reallocated the budget. That is what a test is for.

What none of these figures support is skipping the test. Adopting a published benchmark as your own incrementality rate is the same error as trusting attributed ROAS, differing only in which number you are believing without evidence.

10/12 Section

Turning A Result Into A Decision

Most incrementality work dies here, producing a slide instead of a budget change.

If incremental ROAS clears your margin breakeven, the campaign is profitable at its true rate even if the honest number is far below the reported one. Keep it. This is a common and underappreciated outcome, and cutting on the basis that the number came down is an overreaction.

If it sits below breakeven, reduce rather than eliminate first. Branded search in particular has defensive value against competitors bidding on your terms, which is real and does not appear in an incrementality figure.

If lift is indistinguishable from zero, that is a strong signal and you should act on it. Reallocate to something with demonstrated incrementality, which usually means the category and competitor targeting that attributed reporting has been undervaluing all along.

If the result is ambiguous, resist re-running immediately. Each test costs ranking. Improve the design and wait for a period where the read will be cleaner, avoiding seasonal distortion.

Whichever way it goes, the comparison is against your contribution margin rather than against a ROAS target inherited from somewhere. Our contribution margin playbook covers building that per SKU, and revenue against actual profit covers why the distinction changes conclusions.

11/12 Section

Common Ways Tests Fail

Recognisable patterns, each of which invalidates the result rather than merely weakening it.

  1. Improvising the holdout inside live trafficking. Deciding the control group after campaigns are running invites contamination. Build the control logic in advance.
  2. Measuring attributed sales instead of total sales. The most common error. You will observe credit moving and mistake it for demand moving.
  3. Running during a distorted period. Prime Day, holiday peaks, and your own promotions all break the comparison between test and control.
  4. Changing something else mid-test. A price change, a listing update, a stockout. Any of these ends the experiment whether or not anyone notices.
  5. Underpowering it. Too little spend, too few products, too short a window. Produces an inconclusive result at full cost.
  6. Ignoring the ranking effect. Amazon-specific, and it inflates measured incrementality in a direction that happens to flatter the campaign you were testing.
  7. Deciding the threshold afterwards. Choosing what counts as significant once you have seen the number guarantees you find what you wanted.

Number four deserves particular attention on Amazon because stockouts are common and devastating to a test. A product that runs out mid-experiment has left the study, and if it was in your test group the result is not salvageable.

12/12 Section

When Not To Test

Incrementality testing is fashionable and it is not always the right next step. Several situations call for something else.

When your campaign structure has overlap problems. If several campaigns target the same shoppers without exclusions, fix that first. You will resolve most double-counting without spending ranking on an experiment.

When spend is too small to read. Below meaningful volume, no design will separate signal from noise. You will pay the full ranking cost for an inconclusive answer, which is the worst available outcome.

When you would not act on the result. If branded search is going to keep running for defensive reasons whatever the number says, the test is expensive curiosity.

When products are ramping. Never test on products still establishing rank. The ranking cost is highest and the baseline is least stable exactly there.

During Q4. Distorted demand, maximum ranking value, and the highest cost of a mistake. Test in a quiet period.

The Cheaper Thing To Do First

Before designing an experiment, look at what proportion of your ad-attributed revenue comes from exact-match branded terms. If it is large, you already know most of that revenue was probably coming anyway, and you can act on that suspicion by trimming branded bids modestly and watching total revenue rather than by running a formal test. That is not rigorous and it is nearly free, and for many brands the directional answer is enough to move budget sensibly.

Save the formal test for when the stakes justify the ranking cost, meaning a large budget line, a genuine internal disagreement, or a decision that will not be made without evidence. Our guide to attribution tracking covers the measurement groundwork that makes any of this readable, and Amazon DSP for growing brands covers whether the audience-holdout route is available to you at all.

Key Takeaways

What To Remember

  • Attribution is bookkeeping and incrementality is causation. ROAS records which sales were credited to a campaign, never which would have failed to happen without it.
  • Pausing ads on Amazon degrades organic ranking, because ranking responds to sales velocity. That inflates measured incrementality and leaves a real cost after the test ends.
  • Test on a controlled subset of products against a matched set, never the whole catalog, and never on products still establishing rank.
  • Measure total units or revenue, not attributed sales. Watching credit move between campaigns is attribution shift, not an incrementality test.
  • Set the confidence threshold, minimum detectable effect, and both decision paths before launch. Deciding afterwards guarantees you find what you wanted.
  • Circulating benchmarks are vendor-sourced. Branded search figures of 1.5x to 3x incremental are a hypothesis to test on your account rather than a number to adopt.
  • Judge the result against contribution margin breakeven, not against the reported ROAS. A campaign can be worth keeping at a much lower true rate.
Sources

Where This Came From

  1. Amazon Ads, Amazon Marketing Cloud, for the clean-room analysis capability underpinning audience holdout measurement.
  2. Industry and measurement-vendor writing on incrementality methodology, covering audience holdouts, geographic holdouts, product-subset pause tests, and platform lift studies, along with the recommendation to fix confidence thresholds before launch and to size spend to the read required.
  3. Measurement-vendor writing on geographic test prerequisites, including first-party performance data at state level reported weekly or daily and at least six months of clean historical data before testing.
  4. Measurement-vendor benchmark claims discussed critically in section 09: branded search incremental ROAS of roughly 1.5x to 3x against 10x or higher attributed, retargeting of roughly 1.5x to 2.5x, a claim that incremental ROAS across DTC brands runs 40 to 70 percent below platform-reported figures, and a single case study in which a grocery chain found no measurable lift from non-branded paid search across 12 test markets. All originate from companies selling measurement services and none publish full methodology.
  5. Industry writing noting that most brands claiming to test incrementality are measuring attribution shift, and that Amazon tests require patience to weather a short-term ranking dip, which corroborates the velocity mechanism described in section 06.

Questions

Twelve things advertisers ask about incrementality
What is the difference between attribution and incrementality?

Attribution assigns credit for conversions that happened, according to rules about which touchpoints count. Incrementality asks how many of those conversions would have occurred anyway without the advertising. Attribution is bookkeeping performed on observed events; incrementality requires an experiment, because the counterfactual is not present in the data.

Why is branded search the first thing to test?

Because a shopper typing your exact brand name has already chosen you and would likely find your organic listing regardless. Branded campaigns routinely report the highest ROAS in an account while creating the least new demand, which makes them the largest gap between what is credited and what was caused.

What makes incrementality testing harder on Amazon?

Amazon's organic ranking responds to sales velocity, so pausing ads reduces velocity, which degrades organic position, which reduces organic sales further. You end up measuring ads-off plus ranking decay rather than ads-off alone, which inflates apparent incrementality, and the lost ranking persists after the test finishes.

How do I stop the ranking effect ruining my test?

You reduce it rather than eliminate it. Test on a controlled subset of products against a matched set, choose products with established stable rank rather than recent launches, keep the window toward two weeks rather than four, and prefer an audience holdout over a product pause if you have DSP, since that preserves aggregate velocity.

What test designs are available on Amazon?

Four. An audience holdout inside DSP, which is the cleanest randomisation but requires a DSP entity. A geographic holdout comparing matched regions. A pause on a controlled subset of products against a matched set, which is most practical for sellers without DSP. And platform-run lift studies, which are convenient but graded by the party being measured.

How long should an incrementality test run?

Two to four weeks is the typical window. Shorter rarely accumulates enough volume to separate effect from noise. Longer accumulates contamination and, on Amazon specifically, more ranking damage. The shorter end of that range is usually the better compromise given the velocity cost.

What should I measure during the test?

Total units or total revenue across the business for the exposed and withheld groups, not attributed sales. Watching attributed sales move between campaigns is measuring credit relocating rather than demand changing, and it is the most common way brands convince themselves they ran an incrementality test when they did not.

What is a normal incremental ROAS for branded search?

Measurement vendors commonly report figures around 1.5x to 3x against attributed returns of 10x or higher, but those come from companies selling measurement services and do not publish full methodology. Treat them as a hypothesis worth testing on your own account rather than as a benchmark you can adopt without evidence.

What if the test shows no lift at all?

That is a legitimate and useful finding rather than a failed test. One reported case involved a grocery chain pausing non-branded paid search across twelve markets, finding no measurable lift, and reallocating the budget elsewhere. Zero is a possible answer and acting on it is the entire point of testing.

Should I cut a campaign if incremental ROAS is lower than reported?

Not automatically. The right comparison is incremental ROAS against your contribution-margin breakeven, not against the reported figure. A campaign can be genuinely profitable at a much lower true rate. Reduce before eliminating, and remember branded search carries defensive value against competitors that no incrementality number captures.

When should I not run an incrementality test?

When campaign overlap has not been fixed, since that is a cheaper cause of double-counting. When spend is too small to produce a readable signal. When you would not act on the answer. On products still establishing rank. And during Q4, when demand is distorted and ranking is most valuable.

Is there a cheaper alternative to a formal test?

Look at what share of your ad-attributed revenue comes from exact-match branded terms. If it is large, trim branded bids modestly and watch total revenue rather than attributed revenue. That is directional rather than rigorous and it is nearly free, and for many brands it is enough to move budget sensibly without paying the ranking cost.

Ian Smith, founder of Evolve Media Agency
Ian Smith
Founder, Evolve Media Agency

Ian founded Evolve Media Agency in 2017 and has worked in ecommerce since 2015. He has built and sold three companies and generated more than $25M in client revenue through email marketing, and he writes about marketplace strategy, listing optimization, and AI search for ecommerce brands.

Read Ian's Story

Your Best-ROAS Campaign Might Be Your Least Incremental

Before designing an experiment, it is worth checking whether cheaper structural fixes would resolve most of the double-counting. Often they do.

?
The Counterfactual