Group name
Website heatmap
Webflow Optimize
Webflow A/B testing
Voice of customer
User journey map
User behavior analytics
Usability testing
Trust signals
Tree testing
Time on page
Survey design
Statistical significance
Split URL testing
Split testing
Social proof
Session replay
Session replay tools
Session recording
Sequential testing
Segmentation analysis
Scroll map
Scroll depth
Scarcity marketing
Revenue per visitor
Rage click
PIE framework
Novelty effect
Multivariate testing
Mobile conversion rate
Microsoft Clarity
Micro conversion
Message match
Macro conversion
LIFT model
Landing page optimization
Landing page conversion rate
Information scent
ICE score
Hotjar
Holdout group
Hick's law
Heatmap tools
Guardrail metrics
Goal completion
Funnel analysis
Form analytics
Form abandonment
Five second test
Fitts's law
Exit rate
Exit intent popup
Event tracking
Drop-off rate
Dead click
CRO tools
CRO audit
Choosing CRO tools
Conversion funnel
Cohort analysis
Cognitive load
Click map
Checkout optimization
Cart abandonment
Bounce rate
Bayesian A/B testing
Average order value
Attention map
Anchoring bias
Above the fold
A/B testing tools
What is sequential testing?
Sequential testing is a family of methods that evaluate a test repeatedly as data arrives and stop it as soon as the evidence is conclusive, without the inflated false positive rate that repeated checking causes in a fixed-horizon test. It solves the peeking problem by adjusting the decision threshold for the number of looks.
What is the peeking problem?
A fixed-horizon test is valid at one moment: the predetermined sample size. Its 5% false positive rate assumes you look once.
Because results fluctuate, the confidence figure on two identical variants drifts up and down as data accumulates, and on some days it will sit above 95% through chance alone. Each check is another chance to catch it there, so checking daily for two weeks pushes the probability of crossing at some point far above 5%. Stop the moment it crosses and you have manufactured a winner from noise, and how far above 5% your real error rate sits depends on how often you looked.
This is the most common invalidating practice in conversion testing, and it does not feel like cheating. It feels like being attentive. The dashboard updates, someone looks, the number is green, the test ends.
How does sequential testing fix it?
By setting a threshold that accounts for the number of looks in advance.
Two families are used in practice.
Group sequential methods. Define a small number of interim analyses, for example at 25%, 50%, and 75% of the planned sample. Each interim look uses a stricter threshold than 95%, with the thresholds chosen so the overall false positive rate across all looks stays at 5%. Requires committing to the schedule of looks up front.
Always-valid inference. Uses confidence sequences that remain valid at every moment, so you can check continuously. Several modern testing platforms implement this, which is why their dashboards can be watched without the usual warning.
Both work by making early stopping harder than a naive 95% read. An effect that would cross the line on day three under fixed-horizon rules needs to be substantially larger to cross under sequential rules.
What does it cost?
Sequential methods are not free, and the tradeoff is rarely stated.
Larger maximum sample size. A sequential test that runs to completion needs more data than a fixed-horizon test of the same power, because the correction for multiple looks is paid whether or not you stop early. You gain the option to stop early and pay for the option.
Overstated effect sizes when you do stop early. A test ends early because the evidence arrived quickly, and evidence arrives quickly when the early noise happens to favor the variant. The conclusion about which variant won survives that. The lift attached to it is systematically generous, so treat an early stop's number as a ceiling rather than a forecast, and say so when someone builds a revenue projection on it.
Complexity. The thresholds are not something to compute by hand, so this is a decision to use a platform that implements it rather than a technique to apply manually.
When should you use it?
Use it when stopping early has real value. High traffic where a bad variant costs money every day it runs, or a change you need to roll out quickly.
Use it when you cannot stop people looking. If stakeholders will watch the dashboard regardless, a method that stays valid under observation is more honest than a fixed-horizon test everyone peeks at.
Skip it for low traffic. A test that needs six weeks to reach minimum power will not stop early under any method, so the larger maximum sample is pure cost.
Skip it when the constraint is duration, not sample. Tests must run a full business cycle regardless of statistics, typically two weeks, to cover weekday and weekend behavior and to let novelty effects decay. A sequential method reaching significance on day four does not license stopping on day four.
Related terms
Statistical significance · Bayesian A/B testing · Split testing · Novelty effect · Guardrail metrics
Deeper reading: What is structured A/B testing. Service: Conversion Rate Optimization.
FAQ
Is sequential testing the same as Bayesian testing?
No, though both tolerate monitoring. Sequential testing is a frequentist approach that adjusts thresholds for multiple looks. Bayesian testing uses a different inferential framework entirely. Both address peeking, by different means.
Can you stop a sequential test whenever you want?
You can stop whenever the adjusted threshold is met, which is the point of the method. Practical constraints still apply: run a full business cycle, and expect the reported effect size to be optimistic if you stop very early.
Does my testing tool support sequential testing?
Search its documentation for "always valid", "sequential", or "alpha spending". If the tool reports a plain confidence figure that updates live with no mention of an adjustment, treat the dashboard as fixed-horizon and read it once, at the sample size you planned.
What if I already peeked at a fixed-horizon test?
The result is compromised, and how badly depends on how often you looked and whether looking influenced the stopping decision. The clean remedy is to run to the planned sample size and read it once, treating earlier looks as informal.