Group name
Website heatmap
Webflow Optimize
Webflow A/B testing
Voice of customer
User journey map
User behavior analytics
Usability testing
Trust signals
Tree testing
Time on page
Survey design
Statistical significance
Split URL testing
Split testing
Social proof
Session replay
Session replay tools
Session recording
Sequential testing
Segmentation analysis
Scroll map
Scroll depth
Scarcity marketing
Revenue per visitor
Rage click
PIE framework
Novelty effect
Multivariate testing
Mobile conversion rate
Microsoft Clarity
Micro conversion
Message match
Macro conversion
LIFT model
Landing page optimization
Landing page conversion rate
Information scent
ICE score
Hotjar
Holdout group
Hick's law
Heatmap tools
Guardrail metrics
Goal completion
Funnel analysis
Form analytics
Form abandonment
Five second test
Fitts's law
Exit rate
Exit intent popup
Event tracking
Drop-off rate
Dead click
CRO tools
CRO audit
Choosing CRO tools
Conversion funnel
Cohort analysis
Cognitive load
Click map
Checkout optimization
Cart abandonment
Bounce rate
Bayesian A/B testing
Average order value
Attention map
Anchoring bias
Above the fold
A/B testing tools
What is statistical significance?
Statistical significance is the judgment that a difference between two test variants is unlikely to have arisen by chance alone. In conversion testing it is usually expressed as a confidence level, where 95% significance means that if there were genuinely no difference between the variants, a result this extreme would appear less than 5% of the time.
What does 95% significance actually mean?
It is a statement about the behavior of your testing procedure, not about the probability that your variant is better.
Formally, significance answers: assuming the two variants perform identically, how often would random noise produce a difference at least this large? A p-value of 0.05 means 5% of the time. Setting a 95% confidence threshold means accepting that, when two variants are genuinely identical, one test in twenty still declares a winner.
The common misreading is "95% significance means a 95% chance the variant is better." That is a different quantity, and it depends on how likely the variant was to be better before you ran the test. In a program where most variants turn out to be no better than control, a result sitting right at 95% is more often noise than the number suggests.
Key takeaway: significance controls how often you are fooled across many tests. It says less than people think about the specific test in front of you.
What does statistical significance not tell you?
Four things, each of which causes real damage when assumed.
It does not tell you the effect is large. With enough traffic, a 0.1% lift reaches significance. Significant and worth shipping are different questions. Read the confidence interval around the effect, which tells you the range of lifts the data is consistent with, and decide against the bottom of that range rather than the midpoint.
It does not tell you the effect will last. A winning variant may be winning because it is new. See novelty effect.
It does not tell you the result generalizes. A test run over two weeks in December on desktop traffic describes that period, that season, and that device.
It does not validate the metric. A variant that significantly increases clicks on a CTA while significantly decreasing qualified demos is a significant loss. See guardrail metrics.
How do teams accidentally fake significance?
Three practices, all common, all of which inflate false positives well beyond the 5% the threshold promises.
Peeking and stopping. Checking the dashboard daily and stopping when it crosses 95%. Because the result fluctuates, a test on two identical variants drifts above 95% on some days on its own, and daily checking is a policy of stopping on whichever day that happens. Fixed-horizon tests must run to their predetermined sample size. If you want to stop early, use a method built for it. See sequential testing.
Testing many metrics. Twenty metrics checked at 95% each will hand you about one false winner per test, purely from the arithmetic of twenty chances at a one-in-twenty error. Declare one primary metric before starting.
Segment mining after the fact. Finding no overall effect, then discovering the variant won among mobile users in Canada. With enough segments, something always wins. Segments must be specified in advance to count as evidence.
What should you decide before a test starts?
Write these down before traffic is allocated. A test whose parameters were chosen after seeing data is not a test.
- Primary metric. One. The number that decides the test.
- Minimum detectable effect. The smallest lift worth shipping for. This drives sample size, and choosing it honestly usually reveals the test needs more traffic than expected.
- Sample size and duration. Calculated from baseline rate, MDE, and desired power. Run at least one full business cycle, typically two weeks, regardless of what the calculator says, because weekday and weekend traffic differ.
- Guardrail metrics. What must not get worse.
- Stopping rule. The date or sample size at which you look, decide, and stop.
That discipline is the difference between a testing program and a series of anecdotes. It is covered in depth in building data-backed CRO hypotheses.
Related terms
Split testing · Sequential testing · Bayesian A/B testing · Novelty effect · Guardrail metrics
Deeper reading: What is structured A/B testing. Service: Conversion Rate Optimization.
FAQ
Is 95% confidence the right threshold?
It is a convention, not a law. Raise it when a wrong decision is expensive or hard to reverse. Lower it when the change is cheap, easily reverted, and the cost of missing a real improvement exceeds the cost of a false positive. State the choice before the test rather than after.
Can a test be significant and still wrong?
Yes. Significance controls one specific error, the false positive rate under repeated sampling. It offers no protection against a broken implementation, a contaminated sample, a seasonal effect, or a metric that does not represent business value.
What if my test never reaches significance?
That is a result. It usually means the effect is smaller than your minimum detectable effect, which tells you the change is not worth shipping at your traffic level. Report it as inconclusive, keep the simpler variant, and move to a bigger hypothesis.
Do I need statistics to run CRO?
You need enough to avoid the three faking patterns above. Sample size calculators and modern testing platforms handle the arithmetic. The judgment about what to measure, when to stop, and what counts as a win is the part no tool supplies.