Group name

Website heatmap

Behavior analytics

Voice of customer

Research methods

User journey map

Research methods

User behavior analytics

Behavior analytics

Usability testing

Research methods

Trust signals

Page levers

Tree testing

Research methods

Time on page

Metrics and funnel

Survey design

Research methods

Social proof

Page levers

Session replay

Behavior analytics

Session recording

Behavior analytics

Segmentation analysis

Metrics and funnel

Scroll map

Behavior analytics

Scroll depth

Behavior analytics

Revenue per visitor

Metrics and funnel

Rage click

Behavior analytics

PIE framework

Research methods

Mobile conversion rate

Metrics and funnel

Micro conversion

Metrics and funnel

Message match

Page levers

Macro conversion

Metrics and funnel

LIFT model

Research methods

ICE score

Research methods

Hotjar

Tools

Hick's law

Page levers

Goal completion

Metrics and funnel

Funnel analysis

Metrics and funnel

Form analytics

Behavior analytics

Form abandonment

Behavior analytics

Five second test

Research methods

Fitts's law

Page levers

Exit rate

Metrics and funnel

Event tracking

Metrics and funnel

Drop-off rate

Metrics and funnel

Dead click

Behavior analytics

CRO audit

Research methods

Conversion funnel

Metrics and funnel

Cohort analysis

Metrics and funnel

Cognitive load

Page levers

Click map

Behavior analytics

Cart abandonment

Metrics and funnel

Bounce rate

Metrics and funnel

Average order value

Metrics and funnel

Attention map

Behavior analytics

Anchoring bias

Page levers

Above the fold

Page levers
This is some text inside of a div block.
Created:
updated:

What is sequential testing?

Sequential testing is a family of methods that evaluate a test repeatedly as data arrives and stop it as soon as the evidence is conclusive, without the inflated false positive rate that repeated checking causes in a fixed-horizon test. It solves the peeking problem by adjusting the decision threshold for the number of looks.

What is the peeking problem?

A fixed-horizon test is valid at one moment: the predetermined sample size. Its 5% false positive rate assumes you look once.

Because results fluctuate, the confidence figure on two identical variants drifts up and down as data accumulates, and on some days it will sit above 95% through chance alone. Each check is another chance to catch it there, so checking daily for two weeks pushes the probability of crossing at some point far above 5%. Stop the moment it crosses and you have manufactured a winner from noise, and how far above 5% your real error rate sits depends on how often you looked.

This is the most common invalidating practice in conversion testing, and it does not feel like cheating. It feels like being attentive. The dashboard updates, someone looks, the number is green, the test ends.

How does sequential testing fix it?

By setting a threshold that accounts for the number of looks in advance.

Two families are used in practice.

Group sequential methods. Define a small number of interim analyses, for example at 25%, 50%, and 75% of the planned sample. Each interim look uses a stricter threshold than 95%, with the thresholds chosen so the overall false positive rate across all looks stays at 5%. Requires committing to the schedule of looks up front.

Always-valid inference. Uses confidence sequences that remain valid at every moment, so you can check continuously. Several modern testing platforms implement this, which is why their dashboards can be watched without the usual warning.

Both work by making early stopping harder than a naive 95% read. An effect that would cross the line on day three under fixed-horizon rules needs to be substantially larger to cross under sequential rules.

What does it cost?

Sequential methods are not free, and the tradeoff is rarely stated.

Larger maximum sample size. A sequential test that runs to completion needs more data than a fixed-horizon test of the same power, because the correction for multiple looks is paid whether or not you stop early. You gain the option to stop early and pay for the option.

Overstated effect sizes when you do stop early. A test ends early because the evidence arrived quickly, and evidence arrives quickly when the early noise happens to favor the variant. The conclusion about which variant won survives that. The lift attached to it is systematically generous, so treat an early stop's number as a ceiling rather than a forecast, and say so when someone builds a revenue projection on it.

Complexity. The thresholds are not something to compute by hand, so this is a decision to use a platform that implements it rather than a technique to apply manually.

When should you use it?

Use it when stopping early has real value. High traffic where a bad variant costs money every day it runs, or a change you need to roll out quickly.

Use it when you cannot stop people looking. If stakeholders will watch the dashboard regardless, a method that stays valid under observation is more honest than a fixed-horizon test everyone peeks at.

Skip it for low traffic. A test that needs six weeks to reach minimum power will not stop early under any method, so the larger maximum sample is pure cost.

Skip it when the constraint is duration, not sample. Tests must run a full business cycle regardless of statistics, typically two weeks, to cover weekday and weekend behavior and to let novelty effects decay. A sequential method reaching significance on day four does not license stopping on day four.

Related terms

Statistical significance · Bayesian A/B testing · Split testing · Novelty effect · Guardrail metrics

Deeper reading: What is structured A/B testing. Service: Conversion Rate Optimization.

FAQ

Is sequential testing the same as Bayesian testing?

No, though both tolerate monitoring. Sequential testing is a frequentist approach that adjusts thresholds for multiple looks. Bayesian testing uses a different inferential framework entirely. Both address peeking, by different means.

Can you stop a sequential test whenever you want?

You can stop whenever the adjusted threshold is met, which is the point of the method. Practical constraints still apply: run a full business cycle, and expect the reported effect size to be optimistic if you stop very early.

Does my testing tool support sequential testing?

Search its documentation for "always valid", "sequential", or "alpha spending". If the tool reports a plain confidence figure that updates live with no mention of an adjustment, treat the dashboard as fixed-horizon and read it once, at the sample size you planned.

What if I already peeked at a fixed-horizon test?

The result is compromised, and how badly depends on how often you looked and whether looking influenced the stopping decision. The clean remedy is to run to the planned sample size and read it once, treating earlier looks as informal.

Ask AI about this term

Want more revenue from your existing traffic?

We run CRO for Webflow sites — from audit to A/B testing.

Work with us

Work with us