Group name

Website heatmap

Behavior analytics

Voice of customer

Research methods

User journey map

Research methods

User behavior analytics

Behavior analytics

Usability testing

Research methods

Trust signals

Page levers

Tree testing

Research methods

Time on page

Metrics and funnel

Survey design

Research methods

Social proof

Page levers

Session replay

Behavior analytics

Session recording

Behavior analytics

Segmentation analysis

Metrics and funnel

Scroll map

Behavior analytics

Scroll depth

Behavior analytics

Revenue per visitor

Metrics and funnel

Rage click

Behavior analytics

PIE framework

Research methods

Mobile conversion rate

Metrics and funnel

Micro conversion

Metrics and funnel

Message match

Page levers

Macro conversion

Metrics and funnel

LIFT model

Research methods

ICE score

Research methods

Hotjar

Tools

Hick's law

Page levers

Goal completion

Metrics and funnel

Funnel analysis

Metrics and funnel

Form analytics

Behavior analytics

Form abandonment

Behavior analytics

Five second test

Research methods

Fitts's law

Page levers

Exit rate

Metrics and funnel

Event tracking

Metrics and funnel

Drop-off rate

Metrics and funnel

Dead click

Behavior analytics

CRO audit

Research methods

Conversion funnel

Metrics and funnel

Cohort analysis

Metrics and funnel

Cognitive load

Page levers

Click map

Behavior analytics

Cart abandonment

Metrics and funnel

Bounce rate

Metrics and funnel

Average order value

Metrics and funnel

Attention map

Behavior analytics

Anchoring bias

Page levers

Above the fold

Page levers
This is some text inside of a div block.
Created:
updated:

What is a holdout group?

A holdout group is a segment of visitors deliberately excluded from a change after it ships, kept on the previous experience so the difference can be measured over a longer horizon. It answers the question a completed A/B test cannot: did the improvement persist, and does the accumulation of many shipped changes actually add up.

Why keep a holdout after the test already won?

Because a two-week test measures a two-week effect, and most claimed gains are asserted far beyond that window.

Three specific problems a holdout resolves:

Decay. A variant winning on novelty reverts once the novelty fades. A holdout running for a month after ship shows whether the difference is still there. See novelty effect.

Accumulation. A team shipping twelve wins a year, each claiming 5%, is implicitly claiming the site converts 80% better than a year ago. It almost never does. A long-running holdout on the original experience measures the true combined effect, which is usually far smaller than the sum of the parts.

Interaction. Changes that each win individually can conflict when combined. Sequential tests cannot see this. A holdout comparing the full current experience against the old one can.

Key takeaway: individual tests measure individual changes. Only a holdout measures the program.

What only shows up over months?

Decay, accumulation, and interaction are visible within weeks of shipping. Three more effects need a longer clock, and they are the ones that decide whether a testing program is producing value.

Long-horizon business outcomes. For B2B SaaS with a long sales cycle, a two-week test can only measure a proxy such as demo requests. A holdout running for a quarter can be joined to closed revenue.

Seasonality. A change that wins in one period may lose in another. A holdout spanning several months exposes that.

Cumulative degradation. Each change adds a script, a section, a field. Individually invisible, collectively a slower and heavier page. A holdout on a clean original version reveals the drift.

How do you size and run one?

Size it small. 5 to 10% of traffic is typical. The holdout is a reference point, not a second experiment, and every visitor in it is one denied the improvement you believe in.

Assign persistently and randomly. Same requirements as any split. A visitor in the holdout stays in the holdout, or the comparison means nothing.

Run it long. The value is in the horizon. A month at minimum, a quarter for programs where the business outcome is slow.

Decide the read-out date in advance. A holdout checked whenever someone is curious has the same peeking problem as any test.

Use a program-level holdout, not one per change. One group held on the experience as of a fixed date, against which everything shipped since is measured collectively. Per-change holdouts fragment traffic and answer a question the original test already answered.

Exclude it from other experiments. A holdout contaminated by concurrent tests is not a clean baseline.

What are the costs?

Real, and worth stating plainly.

Foregone conversions. If the changes genuinely work, the holdout group converts worse for as long as it runs. That is the price of knowing.

Operational complexity. Assignment must survive deploys, and anyone building pages needs to know a holdout exists so they do not accidentally break it.

Analytical discipline. A holdout nobody reads is pure cost.

For a site with substantial traffic and a testing program making repeated claims, the cost is usually worth paying, because the alternative is a set of compounding numbers nobody has ever verified. For a site running two tests a year, it is not.

Related terms

Split testing · Novelty effect · Guardrail metrics · Statistical significance · Cohort analysis

Deeper reading: What is structured A/B testing. Service: Conversion Rate Optimization.

FAQ

How big should a holdout group be?

5 to 10% of traffic for a program-level holdout. Large enough to detect a meaningful cumulative difference over months, small enough that the foregone conversions stay acceptable.

How long should a holdout run?

A month at minimum. A quarter or longer when the business outcome you care about is slow, which is typical in B2B SaaS where a demo request converts to revenue months later.

Is a holdout group the same as a control group?

Related but not identical. A control group exists during a test to compare against a variant. A holdout persists after the decision has been made and the winner has shipped, to verify the effect over a longer horizon.

Do you need a holdout if your tests are well run?

Well-run tests still only measure short windows and single changes. A holdout answers whether the effects persisted and whether they combined, which no individual test addresses regardless of how carefully it was run.

Ask AI about this term

Want more revenue from your existing traffic?

We run CRO for Webflow sites — from audit to A/B testing.

Work with us

Work with us