Google's HEART Framework: Measuring UX With Product Data

The HEART framework is Google's way of turning "is this experience any good?" into numbers you can actually watch move. It sorts user-experience quality into five categories (Happiness, Engagement, Adoption, Retention, and Task success) and then pairs each one with a Goals-Signals-Metrics process so a vague ambition ("users should love checkout") becomes a specific number on a chart. That's the whole idea. The letters are the easy part; the discipline is in what you choose not to measure.

I'll say the quiet thing out loud first, because it's the mistake I made for years: you are not supposed to fill in all five letters. Most teams treat HEART like a checklist and end up with twenty-five metrics nobody reads. We'll get to how to avoid that. But the definition comes first.

What the five letters actually mean

The heart framework was created in 2010 by Kerry Rodden, Hilary Hutchinson, and Xin Fu on Google's research team, and first published at the ACM CHI conference under the very sober title Measuring the User Experience on a Large Scale. Fifteen-plus years on, it's still the model most product people reach for, and I think that's because it's honest about a hard truth: user experience is plural. There isn't one number.

Here's each dimension, in plain terms.

Happiness is attitude, how people feel about the product. It's the one you can't get from server logs. You collect it by asking: satisfaction surveys, net promoter, a thumbs-up on a search result, a one-question micro-survey after a task.

Engagement is depth of involvement. How often, how much, how intensely. Sessions per week, actions per session, minutes in the editor. Not "did they show up" but "what did they do while they were here."

Adoption is new uptake. How many people started using the product, or a specific feature, in a given window. New accounts that reach first value. Percentage of eligible users who tried the thing you just shipped.

Retention is survival. Of the people who were here before, how many came back? This is the one that quietly decides whether a business exists, and it's the one vanity dashboards are worst at.

Task success is efficiency and effectiveness. Can people actually get the job done? Completion rate, time-on-task, error rate. If someone lands on your checkout and can't finish it, no amount of Happiness survey theater will save you.

Notice these aren't all the same kind of number. Happiness is attitudinal (you have to ask). The other four are behavioral (you can observe them). That split matters, because attitudinal data is expensive and slow and behavioral data is cheap and fast, and teams that forget the difference end up either ignoring how users feel or drowning in survey requests.

The part everyone skips: Goals-Signals-Metrics

A letter on its own is useless. "We care about Engagement," okay, engagement of what, measured how, meaning what? The framework's real engine is the second half, a little procedure called Goals-Signals-Metrics that also came out of Google.

It runs top to bottom, and the order is the point:

  1. Goal. One sentence describing what success looks like for the user, in the category you picked. Not "increase revenue." Something like "shoppers complete checkout without confusion."
  2. Signal. The observable behavior or stated attitude that tells you the goal is happening or failing. What would you see if it were working?
  3. Metric. The specific number, on a specific denominator, you'll track over time.

Think of it like planning a route across a city. The goal is the neighborhood you want to reach. The signals are the landmarks that tell you you're heading the right way. The metric is the actual transit line you take and the arrival time you check your watch against. Skip straight to "let's track a metric" and you're on a train with no idea where it's going, which, funnily enough, describes most analytics dashboards I've inherited.

A worked example: the checkout flow

Let me fill the grid in for one concrete feature, because abstract frameworks are where good intentions go to die. Say you're redesigning a checkout flow for a small e-commerce app. A redesign of a complex form points you toward two dimensions: Task success (can they finish?) and Happiness (did it feel painful?). We'll deliberately leave the other three letters blank. More on why in a second.

HEART dimension Goal Signal Metric
Task success Shoppers complete checkout on the first attempt without errors Reaching the confirmation screen; hitting a validation error; abandoning mid-form Checkout completion rate; median time-to-order; checkout error rate per 100 starts
Happiness The flow feels effortless rather than stressful A post-purchase rating; whether people mention "confusing" or "slow" in feedback One-question CSAT after order ("How easy was that?", 1–5); % of ratings ≤ 2

Two rows. Five metrics total, and honestly you could ship with three. When teams do run this well, the numbers get specific fast. One published checkout example reported completion up 420 basis points, median time-to-order down 19%, and the error rate down 23% after a redesign. Whether or not you hit those exact figures, that's the shape of a Task-success win: more people finishing, faster, with fewer dead ends.

Compare that to the dashboard this feature would have gotten without the framework: total pageviews on the cart, raw add-to-cart count, aggregate session time. All up and to the right, all completely silent on whether checkout got better. Those are the metrics you show an investor, not the ones you fix a product with. (I have built exactly that dashboard. It got a lot of nods in meetings and improved nothing.)

Why you pick two dimensions, not five

So why leave three letters empty? Because measurement isn't free, and I don't only mean the vendor bill.

Every signal you decide to track is an event someone has to instrument, name consistently, keep clean, and then look at. Track all five HEART categories on every feature and you get a wall of twenty-plus numbers where nothing is clearly the one to move. Amplitude's guidance on HEART puts it bluntly: high-performing teams rarely track all five dimensions for a given initiative, and they pick the two or three that match the goal, because optimizing everything at once produces noisy dashboards and weak insight. The Interaction Design Foundation's 2026 write-up adds a scoping rule I lean on: HEART is a feature-level tool. It's the wrong instrument for a single microinteraction (too granular) and for an entire product family (too broad). One feature, one goal, one or two letters.

The selection isn't guesswork. The original paper suggests matching the dimensions to what the work actually is:

What you're doing Dimensions that usually earn their place
Launching a brand-new feature Adoption + Task success
Deepening an existing flow Engagement + Retention
Redesigning a confusing interface Happiness + Task success
Fighting churn Retention + Happiness

These are defaults, not commandments. But starting from two forces the good question, why this number?, instead of the lazy one, what can we track?

There's a reason this restraint pays off beyond tidiness. Userpilot's product-analytics writing in 2025 made the point that teams tracking fewer metrics but reviewing them more often ship faster and decide better, while overloaded dashboards mostly get ignored. A metric nobody opens in a month isn't a metric. It's a cost. If you want the bigger structure that decides which of these HEART numbers ladders up to the outcome that matters, a north-star metric tree is the companion piece. HEART tells you a feature is healthy; the tree tells you whether that health is pointed at the business.

Where this goes wrong

A few failure modes I've watched play out, so you can skip them.

Treating Happiness as free. Attitudinal data needs someone to answer a survey, and survey fatigue is real. If you slap a CSAT prompt on every screen, response rates rot and the people who do answer are the angriest ones. Sample it. Ask rarely, at the moment a task ends, and treat the number as directional rather than precise.

Confusing a metric with a goal. "Increase engagement" is not a goal. It's a metric wearing a goal's coat. More engagement can mean people love your product or that they're lost and clicking around trying to find the exit. Engagement went up the week I once shipped a genuinely worse navigation, because confused users generate a lot of clicks. Without the Goal and Signal rows above it, a metric can't tell you which story it's telling.

Vanity metrics in a nice frame. HEART doesn't immunize you against measuring the easy thing instead of the true thing. Raw MAU, total downloads, and pageviews will happily sit in a "HEART dashboard" and look official. Ask of every metric: if this number doubled overnight, would a real user be better off? If you can't answer yes, it's decoration.

Averages hiding the story. A median time-to-order of nine seconds sounds fine until you notice it's five seconds for returning customers and forty for first-timers. HEART metrics are worth almost nothing un-segmented. New versus returning, device, plan tier: cut the number before you trust it.

A quick FAQ

Is HEART only for big companies like Google? No. It scales down better than most Google-born frameworks. A two-person team shipping one feature can fill in a single row of the grid on a whiteboard in ten minutes. The discipline of Goal → Signal → Metric matters more the smaller you are, because you have less time to waste on numbers that don't act.

HEART vs. AARRR (pirate metrics), which one? Different jobs. AARRR maps a funnel across the whole business lifecycle (acquisition through revenue). HEART measures the quality of a specific experience inside that funnel. They're complementary: pirate metrics tell you people drop off at checkout; HEART, applied to checkout, tells you whether it's a Task-success problem or a Happiness one.

Do I need a special tool for this? Not really. HEART is a way of choosing metrics, not a product you buy. Any analytics setup that can track events and run a light in-app survey covers the behavioral four and Happiness respectively. The framework is upstream of the tooling, so decide the two dimensions first, then check your stack can measure them.

How often should I revisit the grid? When the feature's goal changes. A flow you're launching (Adoption + Task success) becomes, six months later, a flow you're deepening (Engagement + Retention). Same feature, different letters. Re-running GSM at that pivot is cheaper than dragging along metrics that stopped mattering.

That's HEART. Five categories, a two-step process to make each one concrete, and one rule that does most of the work: measure the two things this feature is actually for, and leave the other three letters alone until they earn their spot.