Group name

Website heatmap

Behavior analytics

Voice of customer

Research methods

User journey map

Research methods

User behavior analytics

Behavior analytics

Usability testing

Research methods

Trust signals

Page levers

Tree testing

Research methods

Time on page

Metrics and funnel

Survey design

Research methods

Social proof

Page levers

Session replay

Behavior analytics

Session recording

Behavior analytics

Segmentation analysis

Metrics and funnel

Scroll map

Behavior analytics

Scroll depth

Behavior analytics

Revenue per visitor

Metrics and funnel

Rage click

Behavior analytics

PIE framework

Research methods

Mobile conversion rate

Metrics and funnel

Micro conversion

Metrics and funnel

Message match

Page levers

Macro conversion

Metrics and funnel

LIFT model

Research methods

ICE score

Research methods

Hotjar

Tools

Hick's law

Page levers

Goal completion

Metrics and funnel

Funnel analysis

Metrics and funnel

Form analytics

Behavior analytics

Form abandonment

Behavior analytics

Five second test

Research methods

Fitts's law

Page levers

Exit rate

Metrics and funnel

Event tracking

Metrics and funnel

Drop-off rate

Metrics and funnel

Dead click

Behavior analytics

CRO audit

Research methods

Conversion funnel

Metrics and funnel

Cohort analysis

Metrics and funnel

Cognitive load

Page levers

Click map

Behavior analytics

Cart abandonment

Metrics and funnel

Bounce rate

Metrics and funnel

Average order value

Metrics and funnel

Attention map

Behavior analytics

Anchoring bias

Page levers

Above the fold

Page levers
This is some text inside of a div block.
Created:
updated:

What is an ICE score?

ICE is a prioritization framework that scores each idea on Impact, Confidence, and Ease, usually from 1 to 10, then ranks by the average or product of the three. It exists to make prioritization explicit and arguable rather than driven by whoever spoke last.

How do you score each dimension?

Impact. How much this moves the primary metric if it works. Anchor it to something real: expected lift multiplied by the traffic affected. A 20% improvement on a page 200 people see is a smaller number than a 3% improvement on a page 40,000 people see, and unanchored scoring reliably gets that backwards.

Confidence. How strongly the evidence supports the hypothesis. This is the dimension that separates a testing program from a list of opinions. Score it against evidence type: a finding confirmed by behavioral data and recordings and a prior test scores high, an idea from a competitor teardown scores low.

Ease. How little effort implementation and measurement require. Include the test duration, since an idea needing eight weeks of traffic is not easy regardless of build time.

The output ranks ideas. It does not decide, and treating a 0.3 difference in score as meaningful is over-reading a subjective instrument.

ICE operates on hypotheses that already exist. Producing them is separate work, done by a CRO audit or a structured page review such as the LIFT model. Deciding which page or template deserves the attention in the first place is what the PIE framework is for.

What are the weaknesses?

Scores are subjective and drift. One person's 7 is another's 4, and the same person scores differently in different weeks.

Gaming. Anyone who wants their idea prioritized can score it favorably, and the framework provides no defense.

False precision. Averaging three guesses produces a decimal that looks like measurement.

Ease dominates. Ease is the only one of the three that can be estimated with any accuracy, so it carries the most weight in practice and biases the program toward small changes. The symptom is a year of shipped copy tweaks and no structural test ever attempted.

Impact is the hardest to estimate and matters most, which is an unfortunate combination.

How do you use it without the scores becoming arbitrary?

Define the scale before scoring. Write down what a 3, 5, and 9 mean for each dimension, with examples. Undefined scales are where subjectivity enters.

Anchor impact to arithmetic. Expected lift multiplied by affected traffic, converted to revenue or pipeline. This alone fixes most misranking.

Score confidence against evidence type, using a fixed ladder: prior test result, behavioral data plus qualitative confirmation, behavioral data alone, best practice, opinion.

Score independently, then compare. Two or three people scoring separately, then discussing where they diverge. The discussion is where the value is, and the number is a prompt for it.

Reserve capacity for high-impact work. A fixed share of the testing calendar for structural tests that would never win on ease, which counters the framework's main bias.

Re-score after each test, since confidence changes as evidence accumulates.

Related terms

PIE framework · LIFT model · CRO audit · Statistical significance · Guardrail metrics

Deeper reading: Building data-backed CRO hypotheses. Service: Conversion Rate Optimization.

FAQ

What is the difference between ICE and PIE?

PIE ranks surfaces, ICE ranks hypotheses. PIE scores potential, importance, and ease, weighting traffic value explicitly. ICE scores impact, confidence, and ease, and its confidence dimension is the one PIE has no equivalent for.

Should you multiply or average ICE scores?

Averaging is standard and treats the dimensions as interchangeable. Multiplying punishes a low score on any dimension harder, which better reflects that an idea with no supporting evidence should rank poorly however impactful it might be.

Is ICE scoring worth the effort?

For a team running regular tests, yes, because it makes the reasoning explicit and reviewable. For a team running two tests a year, the discussion matters and the scoring ceremony does not.

Ask AI about this term

Want more revenue from your existing traffic?

We run CRO for Webflow sites — from audit to A/B testing.

Work with us

Work with us