Group name
Website heatmap
Webflow Optimize
Webflow A/B testing
Voice of customer
User journey map
User behavior analytics
Usability testing
Trust signals
Tree testing
Time on page
Survey design
Statistical significance
Split URL testing
Split testing
Social proof
Session replay
Session replay tools
Session recording
Sequential testing
Segmentation analysis
Scroll map
Scroll depth
Scarcity marketing
Revenue per visitor
Rage click
PIE framework
Novelty effect
Multivariate testing
Mobile conversion rate
Microsoft Clarity
Micro conversion
Message match
Macro conversion
LIFT model
Landing page optimization
Landing page conversion rate
Information scent
ICE score
Hotjar
Holdout group
Hick's law
Heatmap tools
Guardrail metrics
Goal completion
Funnel analysis
Form analytics
Form abandonment
Five second test
Fitts's law
Exit rate
Exit intent popup
Event tracking
Drop-off rate
Dead click
CRO tools
CRO audit
Choosing CRO tools
Conversion funnel
Cohort analysis
Cognitive load
Click map
Checkout optimization
Cart abandonment
Bounce rate
Bayesian A/B testing
Average order value
Attention map
Anchoring bias
Above the fold
A/B testing tools
What is an ICE score?
ICE is a prioritization framework that scores each idea on Impact, Confidence, and Ease, usually from 1 to 10, then ranks by the average or product of the three. It exists to make prioritization explicit and arguable rather than driven by whoever spoke last.
How do you score each dimension?
Impact. How much this moves the primary metric if it works. Anchor it to something real: expected lift multiplied by the traffic affected. A 20% improvement on a page 200 people see is a smaller number than a 3% improvement on a page 40,000 people see, and unanchored scoring reliably gets that backwards.
Confidence. How strongly the evidence supports the hypothesis. This is the dimension that separates a testing program from a list of opinions. Score it against evidence type: a finding confirmed by behavioral data and recordings and a prior test scores high, an idea from a competitor teardown scores low.
Ease. How little effort implementation and measurement require. Include the test duration, since an idea needing eight weeks of traffic is not easy regardless of build time.
The output ranks ideas. It does not decide, and treating a 0.3 difference in score as meaningful is over-reading a subjective instrument.
ICE operates on hypotheses that already exist. Producing them is separate work, done by a CRO audit or a structured page review such as the LIFT model. Deciding which page or template deserves the attention in the first place is what the PIE framework is for.
What are the weaknesses?
Scores are subjective and drift. One person's 7 is another's 4, and the same person scores differently in different weeks.
Gaming. Anyone who wants their idea prioritized can score it favorably, and the framework provides no defense.
False precision. Averaging three guesses produces a decimal that looks like measurement.
Ease dominates. Ease is the only one of the three that can be estimated with any accuracy, so it carries the most weight in practice and biases the program toward small changes. The symptom is a year of shipped copy tweaks and no structural test ever attempted.
Impact is the hardest to estimate and matters most, which is an unfortunate combination.
How do you use it without the scores becoming arbitrary?
Define the scale before scoring. Write down what a 3, 5, and 9 mean for each dimension, with examples. Undefined scales are where subjectivity enters.
Anchor impact to arithmetic. Expected lift multiplied by affected traffic, converted to revenue or pipeline. This alone fixes most misranking.
Score confidence against evidence type, using a fixed ladder: prior test result, behavioral data plus qualitative confirmation, behavioral data alone, best practice, opinion.
Score independently, then compare. Two or three people scoring separately, then discussing where they diverge. The discussion is where the value is, and the number is a prompt for it.
Reserve capacity for high-impact work. A fixed share of the testing calendar for structural tests that would never win on ease, which counters the framework's main bias.
Re-score after each test, since confidence changes as evidence accumulates.
Related terms
PIE framework · LIFT model · CRO audit · Statistical significance · Guardrail metrics
Deeper reading: Building data-backed CRO hypotheses. Service: Conversion Rate Optimization.
FAQ
What is the difference between ICE and PIE?
PIE ranks surfaces, ICE ranks hypotheses. PIE scores potential, importance, and ease, weighting traffic value explicitly. ICE scores impact, confidence, and ease, and its confidence dimension is the one PIE has no equivalent for.
Should you multiply or average ICE scores?
Averaging is standard and treats the dimensions as interchangeable. Multiplying punishes a low score on any dimension harder, which better reflects that an idea with no supporting evidence should rank poorly however impactful it might be.
Is ICE scoring worth the effort?
For a team running regular tests, yes, because it makes the reasoning explicit and reviewable. For a team running two tests a year, the discussion matters and the scoring ceremony does not.