Why a scored backlog still stalls
A planning meeting can stall even after every idea has a score. The growth lead wants reactivation emails, product wants a faster mobile flow, and demand generation wants new integration pages. Each proposal sounds plausible, but the team still has one sprint and finite engineering capacity.
Scoring the same backlog with both ICE and RICE shows why the disagreement persists. ICE rewards a team's view of value, certainty, and speed. RICE adds the size of the affected audience and divides by delivery effort. Neither formula removes judgment. Each makes a different judgment visible, and a judgment on the table can be challenged in a way the loudest sponsor's preference cannot.
ICE turns judgment into a fast sort
ICE, popularized by Sean Ellis in growth teams, is a prioritization heuristic defined as:
ICE = Impact x Confidence x Ease
For a compact backlog, use a shared 1-10 scale for all three inputs. Impact estimates movement in a chosen business outcome, such as activated accounts or qualified conversion. Confidence rates the evidence behind that estimate. Ease rates how readily the team can ship and operate the change. Higher is better in every field.
An onboarding checklist might receive Impact 8 because it removes a known setup step, Confidence 8 because session recordings and support tickets point to that friction, and Ease 7 because it uses existing product components. Its ICE score is 8 x 8 x 7 = 448.
That number is a sorting device, not a measurement of economic value. A score of 448 does not mean the checklist is worth 1.24 times as much as an item scored 360. Some teams average the three inputs instead of multiplying them. Multiplication punishes a single weak input harder, so pick one convention and keep it; in the backlog below, averaging keeps the same top three but ties the invite prompt and the mobile fix at 5.67, a tie that multiplication breaks (245 versus 135). The multiplication creates enough separation to force a choice, while retaining a scorecard a team can update in a short planning session.
ICE works best when candidates affect similar audiences, sit in one product area, and have comparable delivery shapes. Under those conditions, explicitly estimating reach may add ceremony without changing the order. It becomes weak when one item affects hundreds of users and another affects nearly every active account.
RICE prices reach before effort
RICE comes from Intercom, where Sean McBride described it in a post on the company blog. Its formula is:
RICE = (Reach x Impact x Confidence) / Effort
Reach is the count of people affected during a fixed time window. A team might define it as eligible users per quarter, activated workspaces per month, or leads entering a campaign during a launch period. The unit and window must stay fixed for every candidate in the comparison.
Intercom's version scores Impact on a fixed relative scale: 3 for massive, 2 for high, 1 for medium, 0.5 for low, and 0.25 for minimal. Confidence is a percentage with three tiers: 100% for high, 80% for medium, 50% for low, and anything below that counts as a moonshot. In the formula it enters as a decimal, 0.8 for 80%. Some teams add intermediate steps such as 0.6 or 0.9, as the example below does; that only works if each step has a written evidence rule. Effort is estimated in person-months across design, engineering, analysis, quality assurance, and rollout work. A lower effort estimate raises the score.
RICE looks more disciplined because reach has a real unit, but its output is only as credible as its inputs. Reach claims depend on clean event definitions and stable audience counts. If tracking cannot distinguish eligible users from users who merely saw a prompt, start with an event taxonomy audit before debating decimal points in a roadmap score.
One backlog produces two orders
The table uses five hypothetical items for a B2B SaaS product. ICE uses 1-10 values. RICE reach means eligible users per quarter, Impact follows the 0.25-3 scale, Confidence is a decimal, and Effort is person-months. These figures are illustrative, not performance forecasts.
| Hypothetical backlog item | ICE I/C/E | ICE score and rank | RICE Reach | RICE I/C/E | RICE score and rank |
|---|---|---|---|---|---|
| Add an onboarding checklist | 8 / 8 / 7 | 448, #1 | 1,200 | 2 / 0.8 / 2 | 960, #3 |
| Publish integration landing pages | 6 / 6 / 9 | 324, #3 | 3,000 | 0.5 / 0.6 / 0.5 | 1,800, #1 |
| Prompt users to invite a teammate | 7 / 7 / 5 | 245, #4 | 400 | 2 / 0.7 / 1 | 560, #5 |
| Improve mobile load time | 9 / 5 / 3 | 135, #5 | 5,000 | 1 / 0.5 / 4 | 625, #4 |
| Send a reactivation email sequence | 5 / 9 / 8 | 360, #2 | 900 | 1 / 0.9 / 0.5 | 1,620, #2 |
Take the checklist calculation. Its ICE score is 8 x 8 x 7 = 448. Its RICE score is (1,200 x 2 x 0.8) / 2 = 960. Both numbers are correct. ICE never asks how many people the checklist touches, and RICE asks little else before dividing by effort.
Integration pages rise from third to first under RICE because they can reach a larger defined audience at low delivery cost. The invite prompt falls because its high expected impact touches fewer eligible users. Mobile performance affects the largest audience, yet four person-months of work and 0.5 confidence leave it fourth, behind three items that need two person-months or less. This is the practical value of scoring both ways: it exposes the assumption that each item deserves equal audience weight.
Ranking gaps reveal hidden assumptions
A large gap between ICE and RICE is a discussion prompt, not a reason to average two scores. Ask which assumption created the gap.
ICE gives the onboarding checklist first place because the team sees a credible, easy change with a strong effect on activation. RICE puts it third: integration pages reach a larger audience, and both the pages and the reactivation sequence cost a quarter of its effort (0.5 versus 2 person-months). Neither result is wrong until the team states its goal. If the quarter's constraint is early-life activation, the checklist may still deserve first place. If the goal is qualified pipeline from a large existing search demand, the pages may be the better bet.
Use ICE when the decision is local: one squad, one audience segment, a short delivery horizon, and little variation in reach. It is also useful for triaging raw ideas before anyone spends time modeling uncertain audience counts. Use RICE when a shared roadmap must compare product, lifecycle, and acquisition work, or when leadership needs to see the opportunity cost of funding a high-impact item with narrow reach.
Confidence deserves special scrutiny for channel proposals. Clicks and form fills can support a hypothesis, but they do not establish incremental business impact. Teams comparing acquisition ideas can borrow the approach used to assess an incrementality claim: record the causal claim, the evidence behind it, and what result would weaken it.
Failure modes that look objective
A spreadsheet can hide disagreement behind neat decimals. Four failure patterns cause false precision.
Confidence inflation. Sponsors often treat confidence as enthusiasm for a project. The mechanism is predictable: people who own a proposal remember favorable anecdotes and discount missing evidence. Spot it when nearly every idea receives 0.8 or 0.9 confidence, even though few have user research, historical results, or a validated causal path. Reserve high confidence for evidence the group can name.
Mixed reach units. One item gets monthly visitors, another gets annual accounts, and a third gets impressions. Multiplication lets the largest raw number dominate the backlog even if the audiences are not comparable. Spot it when the reach column cannot state one unit and one time window in its header. Convert every candidate to the same eligible population before calculating RICE.
Build-only effort estimates. A developer estimates two weeks, while design, experiment instrumentation, legal review, enablement, and monitoring disappear from the denominator. The resulting score favors work that is cheap to code but expensive to make usable. Spot it when effort was supplied by one function or when rollout work appears as an unscored follow-up task.
Overlapping backlog items. A new onboarding flow and an onboarding checklist may compete for the same activation gain. Ranking them independently counts one opportunity twice and creates a misleading portfolio. Spot it when two ideas use the same audience, metric, and causal story. Merge them, or define mutually exclusive variants before scoring.
Calibrate scales before the planning meeting
Calibration means agreeing on what each score means before proposals compete. Nobody has to agree on how the quarter will turn out; the group only has to treat uncertainty and cost the same way for every proposal.
Start with a one-page scoring contract. Define the primary outcome, the reach window, the Impact anchors, the Confidence evidence tiers, and the effort roles included in an estimate. For example, a Confidence value of 0.8 might require prior test results from a comparable segment, while 0.5 could mean a plausible hypothesis supported only by qualitative evidence. Do not call either level a probability unless the team has checked past forecasts against outcomes.
Then use this operating sequence:
- Score independently. Product, growth, engineering, and data partners enter their values before discussion. Early public estimates anchor the room and make later agreement look stronger than it is.
- Discuss the widest spread. If one person assigns 0.9 confidence and another assigns 0.4, ask what evidence each person used. The gap identifies an unknown worth resolving.
- Select a score and save the rationale. A short note such as "reach based on activated workspaces in the prior quarter" makes the next review faster and reveals later unit changes.
- Compare expected and observed outcomes. After release, record actual reach, delivery effort, and the outcome movement tied to the original metric. This will not validate an exact RICE number, but it exposes recurring optimism or weak estimates.
The discipline behind a product interview rubric that reduces bias applies here: named anchors, independent ratings, and a record of disagreement prevent a single persuasive voice from rewriting the scale mid-meeting. For launch work, tie each winning item to an evidence gate it must pass before release, so a high score does not stand in for readiness.
When the extra estimation pays off
ICE buys speed. It is worth the loss of detail when a team is sorting a local backlog with similar audiences and needs a usable order today. RICE buys a clearer cross-functional trade-off: wider reach and total delivery cost are visible instead of implied.
That clarity costs estimation time and can create a false sense of certainty. Pay for RICE when reach or effort could change the funding decision. Otherwise, use ICE, document the assumptions, and revisit the score when evidence changes.