Incrementality testing measures the conversions, revenue, or other business results caused by advertising. It separates those results from outcomes that probably would have occurred without the campaign.
The method creates a counterfactual by comparing a treatment group affected by advertising with a control group that represents expected performance without that advertising. The difference between the groups estimates the campaign’s incremental impact.
This addresses a limitation of traditional attribution. An advertising platform may claim a conversion because a customer viewed or clicked an advertisement before purchasing. That interaction does not prove the advertisement changed the customer’s behavior. Incrementality testing focuses on causation instead of assigning credit based only on recorded touchpoints.
Reliable tests require an appropriate experimental design, sufficient statistical power, contamination controls, and a predetermined analysis plan. When those requirements are met, incrementality testing can help a business determine which channels create additional demand and which ones mainly capture demand that already exists.
Key Takeaways
- Incrementality testing estimates the business results caused by marketing activity.
- A treatment group receives the advertising intervention, while a control group represents expected performance without it.
- Audience holdouts work well when users can be randomly assigned and advertising can be suppressed.
- Geo experiments help measure digital and offline channels without reconstructing individual customer journeys.
- Relative lift, incrementality percentage, and incremental conversions are different metrics with different formulas.
- Statistical power and minimum detectable effect should be evaluated before launching a test.
- A result can be positive, negative, or inconclusive.
- Incrementality results can validate multi-touch attribution and calibrate media mix models.
- A recurring testing roadmap provides more value than a one-time experiment.
What Does Incrementality Testing Measure?
Every observed conversion can be thought of as either baseline or incremental.
Baseline Conversions
Baseline conversions are outcomes expected to occur without the advertising being tested.
They may come from:
- Existing brand awareness
- Organic search
- Repeat customers
- Referrals
- Direct traffic
- Distribution
- Customer loyalty
- Existing purchase intent
- Other advertising campaigns
Incremental Conversions
Incremental conversions are additional outcomes caused by the advertising intervention.
The central question is:
What would have happened if this campaign, channel, or spending increase had not occurred?
Because a business cannot observe the same customer in two realities at the same time, it must estimate the missing outcome using a control group.
That estimated alternative outcome is called the counterfactual.
Why Attribution and Incrementality Produce Different Answers
Attribution assigns credit to interactions recorded before a conversion. Incrementality estimates if marketing activity changed the outcome.
Consider a customer who:
- Visits an e-commerce website.
- Adds a product to their cart.
- Leaves without purchasing.
- Sees a retargeting advertisement.
- Returns and completes the order.
A last-click model may give the retargeting advertisement full credit. A multi-touch model may divide credit among the website visit, product interaction, and retargeting advertisement.
Neither model establishes that the customer needed the advertisement to complete the purchase. The customer may have planned to return already.
Incrementality testing estimates how many similar customers purchase when they are eligible to receive the retargeting campaign compared with customers placed in a holdout group.
Attribution asks:
Which recorded interactions should receive credit?
Incrementality asks:
Did the marketing activity create additional results?
Both methods can be useful, but they support different decisions.
Incrementality Testing vs. A/B Testing
A/B testing compares versions of a marketing asset or experience.
It might ask:
- Which headline generates more clicks?
- Which landing page produces more leads?
- Which creative earns a higher conversion rate?
- Which offer creates more purchases?
Incrementality testing evaluates the effect of running the marketing activity itself.
It might ask:
- Did paid social create additional sales?
- Did branded search capture existing demand or generate new revenue?
- Did connected TV increase purchases?
- Did retargeting create conversions that would not have happened otherwise?
- Did increasing a channel’s budget produce profitable growth?
A/B testing usually helps optimize a tactic that the business has already decided to run. Incrementality testing helps determine the tactic’s causal business value.
All incrementality tests are experiments, but not all marketing experiments measure incrementality.
Important Incrementality Terms
Treatment Group
The treatment group receives the advertising intervention being tested.
Control Group
The control group does not receive the intervention or maintains the normal level of activity.
Holdout
A holdout is the audience, household, account, store, or geographic region intentionally excluded from the advertising treatment.
Counterfactual
The counterfactual estimates what would have happened to the treatment group without the advertising intervention.
Experimental Unit
The experimental unit is what gets assigned to treatment or control. It could be a user, household, company account, store, ZIP code, designated market area, state, or another region.
Incremental Conversion
An incremental conversion is an additional conversion estimated to have been caused by the intervention.
Incremental Revenue
Incremental revenue is the additional revenue attributed to the causal effect of the intervention.
Incremental ROAS
Incremental return on ad spend, or iROAS, compares incremental revenue with the relevant incremental advertising investment.
Statistical Power
Statistical power is the probability that a test will detect an effect of a specified size when that effect is real.
Minimum Detectable Effect
The minimum detectable effect is the smallest lift the experiment is designed to identify reliably.
Confidence Interval
A confidence interval expresses uncertainty around the estimated effect. Wider intervals indicate less precision.
Contamination
Contamination occurs when control units are exposed to the treatment or another meaningful difference affects one group.
Types of Incrementality Tests
The right test depends on the channel, audience, available data, geographic coverage, and business decision.
Platform Conversion-Lift Study
A platform conversion-lift study randomly assigns eligible users to treatment and control groups. The treatment group can receive the campaign, while the control group is withheld from it.
These studies are commonly available for major digital advertising platforms.
Advantages include:
- User-level randomization
- Platform-controlled delivery
- Easier implementation
- Direct connection to platform campaigns
- Faster setup than some independent studies
Limitations include:
- Platform-specific conversion visibility
- Limited cross-channel measurement
- Dependence on platform match rates
- Limited access to underlying user-level data
- Potential differences between platform conversions and finance-verified results
Platform studies can provide useful causal evidence within the platform’s measured environment. Independent analysis can help determine how those results align with total business performance.
User-Level Audience Holdout
A user-level holdout randomly assigns known users to treatment and control groups.
This method may work for:
- Retargeting
- Direct mail
- Loyalty programs
- Subscription campaigns
- Known-customer advertising
The business must be able to identify users, suppress advertising for the control group, and measure outcomes consistently.
Ghost-Ads or Ghost-Bidding Test
A ghost-ads design identifies people who would have been eligible for an advertisement.
Both treatment and control users pass through similar targeting and auction logic. Treatment users receive the advertisement. Control users do not.
This helps reduce the selection bias that can occur when exposed users are compared with people the platform never intended to reach.
Public-Service-Announcement Test
A public-service-announcement test serves the brand advertisement to the treatment group and a neutral advertisement to the control group.
This can help account for the effect of winning the media placement. It may also reduce discrepancies created when one group enters the advertising auction and the other does not.
Geographic Holdout Test
A geographic holdout test continues advertising in selected treatment regions while reducing or withholding it in control regions.
Possible geographic units include:
- ZIP codes
- Counties
- Designated market areas
- States
- Store trade areas
- Countries
Geo experiments can evaluate channels such as:
- Paid search
- Paid social
- Connected TV
- Linear TV
- Streaming audio
- Podcasts
- Out-of-home advertising
- Retail media
- Direct mail
Because geo tests use aggregate market outcomes, they do not require customer-level journey tracking.
Geographic Heavy-Up Test
A geographic heavy-up test increases advertising investment in selected markets while maintaining normal spending in the comparison markets.
This design can estimate the return from additional investment. It is useful when turning a channel off would create excessive business risk.
Matched-Market Test
A matched-market test pairs treatment and control regions using historical similarities such as:
- Revenue
- Conversion volume
- Seasonality
- Customer composition
- Population
- Media costs
- Growth trends
Matching can improve comparability, but it does not provide all the protections of random assignment. Unobserved differences may remain.
Synthetic-Control Test
A synthetic control uses a weighted combination of untreated regions to estimate how the treatment market would have performed without the intervention.
This can create a stronger counterfactual than selecting a single control region. The method requires dependable historical data and careful analysis.
Time-Based Test
A time-based test changes advertising during selected periods and compares performance with other periods.
This method is easier to implement but vulnerable to:
- Seasonality
- Promotions
- Competitor changes
- Economic events
- Weather
- Inventory
- Day-of-week differences
- Long-term growth trends
Time-based tests should be used carefully, especially when the business outcome changes substantially across periods.
Audience Holdouts vs. Geo Experiments
There is no universal minimum number of geographic markets. Feasibility depends on historical volatility, market size, expected lift, available controls, test duration, spending changes, and the statistical method.
How Do You Calculate Incrementality?
The term “lift” can refer to several related metrics. Each should be labeled clearly.
Assume the treatment group has a 3% conversion rate and the control group has a 2% conversion rate.
Absolute Lift
Absolute lift measures the difference in conversion rates.
The advertising increased the measured conversion rate by one percentage point.
Relative Lift
Relative lift compares the conversion-rate difference with the control conversion rate.
The treatment group converted at a rate 50% higher than the control group.
Incrementality Percentage
Incrementality percentage estimates the share of treatment conversions that were incremental.
Approximately one-third of the treatment conversions were incremental under the experiment.
The other two-thirds represent conversions expected under the baseline control rate.
How Do You Calculate Incremental Conversions?
Raw treatment and control conversion counts should not be compared unless both groups have the same size and exposure structure.
Assume:
- Treatment population: 100,000
- Treatment conversions: 3,000
- Control conversion rate: 2%
First calculate the expected baseline conversions in the treatment group:
100,000×2%=2,000
Then calculate incremental conversions:
3,000−2,000=1,000
The campaign generated an estimated 1,000 incremental conversions.
How Do You Calculate Incremental ROAS?
Incremental ROAS compares incremental revenue with the advertising investment associated with creating that lift.
Audience Holdout Example
Assume:
- Incremental conversions: 1,000
- Average revenue per conversion: $100
- Treatment advertising spend: $50,000
Incremental revenue is:
1,000×$100=$100,000
Incremental ROAS is:
The campaign generated an estimated $2 in incremental revenue for every $1 spent.
Geo Heavy-Up Example
A geo heavy-up test should use the additional spending above the counterfactual level.
Assume:
- Treatment-market spending increase: $40,000
- Estimated incremental revenue: $100,000
This estimates the return from the added investment rather than the channel’s full historical budget.
Cost per Incremental Acquisition
Lead-generation, subscription, and customer-acquisition programs may also calculate:
If the business invests $50,000 and generates 200 incremental customers:
The cost per incremental customer is $250.
This may differ substantially from the platform-reported cost per acquisition because the platform metric can include customers who would have converted without the advertising.
Statistical Power and Minimum Detectable Effect
A test can be designed correctly and still fail to provide a useful answer if it does not have enough power.
Power depends on:
- Baseline conversion rate
- KPI volatility
- Audience or market count
- Expected effect size
- Control-group size
- Test duration
- Advertising investment
- Geographic variation
- Conversion frequency
- Desired confidence level
Before launching, define the minimum effect that would change the business decision.
For example:
A paid-social lift below 5% would not justify changing our annual budget, so the test must be capable of detecting at least a 5% lift with an acceptable level of power.
Smaller expected effects generally require larger samples, longer tests, or greater spending differences.
A power analysis helps determine if the proposed test can answer the business question before money is committed.
A 10-Step Incrementality Testing Framework
1. Define the Business Decision
Begin with a decision the test will support.
Examples include:
- Increase prospecting investment
- Reduce retargeting spending
- Protect branded-search spending
- Add connected TV
- Expand into new markets
- Shift funding from one channel to another
A test without a planned decision may produce interesting information that no one acts on.
2. Write a Testable Hypothesis
A useful hypothesis defines the intervention, outcome, and expected direction.
For example:
Increasing paid-social prospecting spending in the treatment markets will generate enough incremental new-customer revenue to exceed our profitability threshold.
3. Select One Primary KPI
Choose a metric that reflects business value:
- Revenue
- Contribution margin
- New-customer revenue
- Qualified leads
- Closed sales
- Subscriptions
- Store visits
- Customer lifetime value
Secondary metrics may add context, but the primary outcome should be selected before the test begins.
4. Choose the Experimental Unit
Decide what will be assigned to treatment or control:
- User
- Household
- Account
- Store
- ZIP code
- Market
- State
- Country
The statistical analysis should match the level of assignment.
5. Define the Minimum Detectable Effect
Identify the smallest effect that would justify a business response.
A statistically detectable change may still be too small to matter financially.
6. Run a Power Analysis
Use historical performance, expected lift, sample size, KPI volatility, group allocation, and test duration to determine feasibility.
If the proposed test is underpowered, consider:
- Extending the duration
- Increasing the treatment intensity
- Adding experimental units
- Choosing a higher-frequency KPI
- Testing a larger campaign
- Revising the minimum detectable effect
7. Create Treatment and Control Groups
Use random assignments when practical.
For geo studies, use randomized markets, matched markets, synthetic controls, or another documented counterfactual method.
Confirm that the groups tracked similarly during the pre-test period.
8. Audit Contamination and Confounders
Document other factors that could affect one group differently, including:
- Promotions
- Price changes
- Creative changes
- Product launches
- Inventory shortages
- Distribution changes
- Competitor activity
- Website updates
- Email campaigns
- Other paid-media campaigns
- Regional events
- Weather
Keep unplanned changes to a minimum during the test.
9. Run the Full Treatment and Observation Period
Do not end the test because early performance looks strong or weak.
The test should cover:
- The predetermined treatment period
- The normal consideration cycle
- The conversion window
- Any planned post-treatment observation period
Upper-funnel campaigns and high-consideration products may require additional time for delayed conversions to emerge.
10. Analyze, Decide, and Retest
Calculate lift, incremental outcomes, efficiency, and uncertainty.
Then apply the predetermined decision rule.
Retest after major changes in:
- Spending level
- Audience strategy
- Creative
- Product mix
- Pricing
- Market conditions
- Channel execution
Incrementality is not a permanent property of a channel. A campaign can be incremental at one spending level and inefficient at another.
Common Incrementality Testing Mistakes
Comparing Exposed and Unexposed Users Without Randomization
People who receive advertisements may differ from those who do not. Comparing naturally exposed users with unexposed users can reproduce the selection bias the test is supposed to remove.
Ending the Test Early
Early results often fluctuate. Stopping after a favorable result increases the risk of a false conclusion.
Changing Multiple Variables
If budget, creative, targeting, landing pages, and promotions all change during the test, the result cannot be attributed to one intervention.
Ignoring Geographic Spillover
People may travel across market boundaries. Television, radio, and out-of-home exposure can cross geographic lines. National campaigns may also reach control regions.
Using an Incomplete Business Outcome
An advertisement may affect website sales, stores, Amazon, retail partners, phone orders, or later subscription revenue.
Testing only website purchases may miss an important halo effect.
Ignoring Conversion Lag
A test window that ends too soon can understate channels with delayed impact.
Treating an Inconclusive Result as Zero Impact
Failure to detect lift does not prove that lift is zero. The effect may be smaller than the minimum detectable effect, or the test may have been too noisy.
Analyzing at the Wrong Level
If markets are assigned to treatment and control, the analysis must recognize markets as the experimental units. Treating every customer inside those markets as an independent randomized observation can overstate precision.
What Is Intent-to-Treat Analysis?
Intent-to-treat analysis compares units according to their original assignment.
A user assigned to the treatment group remains part of the treatment group even if the platform never serves that user an impression. This preserves the benefit of random assignment and estimates the effect of making the campaign available to the treatment group.
Comparing only people who received an impression with those who did not can introduce bias. Impression delivery is influenced by auctions, eligibility, browsing behavior, and platform optimization.
How to Interpret Incrementality Results
Positive and Precise
The estimated effect is positive and the uncertainty range supports a meaningful business impact.
The campaign may be a candidate for increased investment, followed by another test at the higher spending level.
Positive but Inconclusive
The point estimate is positive, but the confidence interval is wide or includes no effect.
The test did not provide enough evidence for a major budget decision. A longer or better-powered test may be appropriate.
No Detectable Lift
The experiment did not detect an effect large enough to distinguish from normal variation.
Possible explanations include:
- The campaign created little incremental value.
- The test was underpowered.
- The treatment was too weak.
- The effect occurred outside the observation window.
- Contamination reduced the difference between groups.
- The selected KPI missed cross-channel or offline effects.
Negative Lift
The treatment group performed worse than the control group.
Investigate:
- Advertisement fatigue
- Poor targeting
- Offer confusion
- Channel conflict
- Website problems
- Inventory limitations
- Random variation
- Test contamination
Do not generalize one negative test to every future campaign on the channel.
Building an Annual Incrementality Testing Roadmap
Incrementality testing provides greater value as a recurring program.
Google recommends planning tests across the year so experiments align with media budgets and promotions and do not overlap in ways that complicate interpretation.[¹]
A testing roadmap might prioritize:
- High-spend campaigns with uncertain incremental value
- Retargeting and branded-search campaigns
- Upper-funnel channels that attribution may undervalue
- New channels
- Budget increases
- Cross-channel halo effects
- Media mix model calibration
- Major seasonal investments
Document each test in a central registry containing:
- Hypothesis
- Owner
- Channel
- Primary KPI
- Test dates
- Experimental unit
- Minimum detectable effect
- Planned analysis
- Result
- Confidence interval
- Budget decision
- Retest date
How Incrementality Supports MTA and MMM
Multi-touch attribution, media mix modeling, and incrementality testing provide different measurement perspectives.
Multi-Touch Attribution
MTA assigns credit across recorded customer interactions. It helps marketers analyze paths and optimize campaigns.
Media Mix Modeling
MMM estimates channel contribution, response curves, diminishing returns, and budget scenarios using aggregate data.
Incrementality Testing
Incrementality testing creates a counterfactual through experimental treatment and control groups. It can validate attributed channel performance and calibrate MMM estimates.
For example, MTA may give branded search substantial credit because it frequently appears before conversion. An experiment may find that many of those customers purchase without the advertisements. The business can then adjust its interpretation of branded-search attribution.
Triple Whale similarly recommends using controlled experiments to validate MTA and MMM conclusions and reconcile conflicting channel results.[²]
Build an Incrementality Program With National Positions
National Positions helps businesses determine which marketing investments create incremental growth through audience holdouts, geographic experiments, budget-expansion tests, multi-touch attribution, and media mix modeling.
The process begins with the decision the test must support. The team then evaluates:
- Available experimental units
- Historical KPI volatility
- Geographic coverage
- Audience size
- Conversion cycle
- Expected lift
- Channel overlap
- Contamination risk
- Data availability
- Business-outcome tracking
A test-readiness assessment can include:
- A testable-channel inventory
- KPI and source-of-truth review
- Audience- or geo-holdout feasibility
- Historical baseline analysis
- Minimum-detectable-effect planning
- Power analysis
- Contamination-risk review
- Recommended test duration
- Metric definitions
- Decision rules
- Post-test budget recommendations
National Positions also integrates incrementality testing with AdBeacon attribution and Meridian media mix modeling to connect experimental evidence with cross-channel measurement and budget planning.[³]
If platform-reported conversions do not align with total revenue or your team cannot determine which channels create additional demand, contact National Positions to request an incrementality test-readiness assessment.
Frequently Asked Questions
What is incrementality testing in simple terms?
Incrementality testing compares a group affected by advertising with a group that represents expected performance without it. The difference estimates the additional results caused by the advertising.
What is the difference between attribution and incrementality?
Attribution assigns conversion credit to recorded marketing interactions. Incrementality testing estimates how many outcomes occurred because of the marketing activity.
What is the difference between incrementality testing and A/B testing?
A/B testing commonly compares versions of creative, landing pages, or offers. Incrementality testing evaluates if running the marketing activity creates additional business outcomes.
What is a holdout test?
A holdout test excludes a selected group from the advertising intervention. Performance in the holdout group helps estimate the baseline outcome.
What is a geo experiment?
A geo experiment changes advertising in selected markets and compares their performance with randomized, matched, or synthetic-control markets.
How many markets does a geo experiment need?
There is no universal minimum. The required number depends on market size, KPI volatility, expected lift, treatment intensity, test duration, and the statistical method. A power analysis should determine feasibility.
How long should an incrementality test run?
The test should run long enough to reach the planned statistical power and capture the normal conversion cycle. The appropriate duration varies by channel, product, audience, and KPI.
What does an inconclusive incrementality test mean?
An inconclusive result means the test did not provide enough evidence to confirm a meaningful positive or negative effect. It does not automatically prove that the campaign has no value.
Can a small business run an incrementality test?
A small business may be able to run a platform conversion-lift study or an audience holdout if it has sufficient conversion volume. Geo experiments generally require enough market-level data and variation to distinguish lift from normal fluctuations.
What is incremental ROAS?
Incremental ROAS divides incremental revenue by the relevant advertising investment. It measures return based on revenue estimated to have been caused by the campaign.
How often should incrementality be retested?
Retest after material changes in budget, targeting, creative strategy, product mix, conversion behavior, or market conditions. High-spend channels may warrant recurring testing.




