Group name

Website heatmap

Behavior analytics

Voice of customer

Research methods

User journey map

Research methods

User behavior analytics

Behavior analytics

Usability testing

Research methods

Trust signals

Page levers

Tree testing

Research methods

Time on page

Metrics and funnel

Survey design

Research methods

Social proof

Page levers

Session replay

Behavior analytics

Session recording

Behavior analytics

Segmentation analysis

Metrics and funnel

Scroll map

Behavior analytics

Scroll depth

Behavior analytics

Revenue per visitor

Metrics and funnel

Rage click

Behavior analytics

PIE framework

Research methods

Mobile conversion rate

Metrics and funnel

Micro conversion

Metrics and funnel

Message match

Page levers

Macro conversion

Metrics and funnel

LIFT model

Research methods

ICE score

Research methods

Hotjar

Tools

Hick's law

Page levers

Goal completion

Metrics and funnel

Funnel analysis

Metrics and funnel

Form analytics

Behavior analytics

Form abandonment

Behavior analytics

Five second test

Research methods

Fitts's law

Page levers

Exit rate

Metrics and funnel

Event tracking

Metrics and funnel

Drop-off rate

Metrics and funnel

Dead click

Behavior analytics

CRO audit

Research methods

Conversion funnel

Metrics and funnel

Cohort analysis

Metrics and funnel

Cognitive load

Page levers

Click map

Behavior analytics

Cart abandonment

Metrics and funnel

Bounce rate

Metrics and funnel

Average order value

Metrics and funnel

Attention map

Behavior analytics

Anchoring bias

Page levers

Above the fold

Page levers
This is some text inside of a div block.
Created:
updated:

What is split testing?

Split testing divides live traffic between two or more versions of a page and measures which produces more conversions. Visitors are assigned randomly and consistently, so each person sees one version for the duration of the test, and the difference in outcome is attributed to the difference in the versions.

Is split testing the same as A/B testing?

In everyday use, yes. The terms are interchangeable and most practitioners treat them as synonyms.

Where a distinction is drawn, it concerns delivery. A/B testing usually means one URL where a script swaps elements client-side. Split URL testing means separate URLs with traffic redirected between them, which suits larger changes. That narrower sense is covered in split URL testing.

Treat "split testing" as the general practice and rely on context.

What makes a split test valid?

Five properties. Missing any one of them means the number at the end is not measuring what you think.

Random assignment. Assignment is decided by chance alone, independent of anything about the visitor. Assignment by day, by device, or by anything correlated with behavior produces a comparison of populations rather than of variants.

Persistent assignment. A returning visitor sees the same variant. Assignment that resets on each visit mixes exposures within one person and destroys attribution.

Concurrent exposure. Both variants run in the same window. Running A this week and B next week compares weeks, not variants, and it passes review because it looks orderly.

A predetermined stopping point. Sample size and duration fixed before launch. See statistical significance for why stopping when it looks good manufactures false winners.

One primary metric. Declared in advance.

Key takeaway: the randomization and the stopping rule carry the validity. The rest is implementation.

What breaks a split test?

Traffic contamination. Your own team, QA sessions, and bots landing in the sample. Filter internal traffic and the staging domain.

Flicker. Client-side tests render the original then rewrite it, producing a visible flash. Visitors seeing the flash behave differently, which is a variable you did not intend to test.

Sample ratio mismatch. A 50/50 split that keeps delivering 53/47 across tens of thousands of sessions is reporting a bug in assignment or tracking, not a coincidence. Check the ratio first, because every other number in the report is computed from it.

Cross-device identity. A visitor on mobile and then desktop is two visitors to most tools, and may see both variants.

A shifting traffic mix. A campaign, a press mention, or a seasonal spike landing mid-test changes who is in the sample. Random assignment keeps the comparison itself fair, but the answer now describes that unusual audience, and carrying it over to normal traffic is an assumption rather than a finding.

Changing the test mid-flight. Editing a variant after launch restarts the experiment, whether or not the tool says so.

How does split testing work on Webflow?

Assignment and reporting come from one of three places: Webflow Optimize natively, an external platform loaded through custom code, or two published pages with traffic split at the source. The full comparison, including which platform limits bite where, is in Webflow A/B testing.

Two Webflow behaviors threaten validity rather than convenience. Republishing the site during a test can alter a live variant mid-flight, and on a site with several editors that publish will come from someone who did not know a test was running. Sessions on the .webflow.io staging domain are your own team, so leaving them in the sample contaminates it with the people who built the variant.

Related terms

Split URL testing · Multivariate testing · Statistical significance · Holdout group · Webflow A/B testing

Deeper reading: What is structured A/B testing. Service: Conversion Rate Optimization.

FAQ

How long should a split test run?

Until it reaches its predetermined sample size, and at minimum one full business cycle, usually two weeks. Weekday and weekend visitors differ, so a test covering only weekdays describes only weekdays.

How much traffic do you need to split test?

Enough conversions, not enough visitors. A page with 50,000 monthly sessions converting at 0.2% has 100 conversions a month, which powers almost nothing. Calculate from your conversion count and the smallest lift worth detecting. Where that calculation says the sample is out of reach, the move is to test bigger: a 40% difference resolves on a fraction of the sample a 2% difference needs, which is why low traffic sites test whole pages and whole offers rather than button colors.

What if both variants perform the same?

Ship the simpler one and treat the hypothesis as unsupported. A flat result at adequate power is real information: that lever does not move this audience.

Ask AI about this term

Want more revenue from your existing traffic?

We run CRO for Webflow sites — from audit to A/B testing.

Work with us

Work with us