The Sean Ellis Test: Measuring PMF with the 40% Rule

The short version: ask your active users "How would you feel if you could no longer use [product]?" with three options, and if at least 40% pick "very disappointed," you have a workable signal of product-market fit. That's the whole test. The number came from Sean Ellis, who ran growth at Dropbox, LogMeIn and Eventbrite before benchmarking roughly a hundred startups and noticing that the ones clearing 40% almost always went on to grow, while the ones below it almost always stalled. Learning Loop's writeup of the Sean Ellis score tells the same story.

So it's one question. The trap is that "one question" makes people careless about the three things that actually decide whether the answer means anything: the exact wording, who you ask, and how many. Get those wrong and you'll get a clean-looking percentage that's quietly garbage.

The question, word for word

Use this wording and don't improvise it:

How would you feel if you could no longer use [product]?

  • Very disappointed
  • Somewhat disappointed
  • Not disappointed (it isn't really that useful)

The score is just the top box: people who answered "very disappointed" divided by everyone who answered, times 100. If 62 of 150 respondents said very disappointed, that's 41%. Above the line.

Why the counterfactual framing ("could no longer use") instead of "how much do you like it"? Because liking is cheap and loss is not. Asking someone to imagine the product taken away forces them to price the gap it would leave. People who'd shrug and switch to a competitor land in "somewhat." Only the ones with no acceptable substitute reach for "very disappointed." That distinction is the entire point, and it's why swapping in a satisfaction scale or an NPS-style 0-to-10 quietly breaks the instrument. They measure sentiment. This measures need.

One more wording note. Keep the third option's parenthetical ("it isn't really that useful"). It gives people permission to admit the product doesn't matter to them, which reduces the polite over-rating you get when the softest option is a bare "not disappointed."

Who you're allowed to ask

Here's where most bad PMF scores are born. The test only means something if you survey people who have actually experienced the product. Fire it at everyone who ever signed up and you're mostly polling ghosts — trialists who bounced on day one, tire-kickers, folks who forgot they registered. They'll mostly answer "not disappointed," because of course they would, and you'll conclude you have no fit when what you actually have is a bad sample.

The rule I use: survey only activated users, recently. CRV's PMF survey guide puts the bar at meaningful engagement, typically two or more weeks of active use or completing the core workflow at least twice. That's a good default. "Activated" should map to your own activation definition, not a generic one. If your product's aha moment is inviting a teammate, then "used it twice" isn't activation and your filter should say so.

And keep it recent. A user who was active four months ago and hasn't returned is answering about a memory. I'd cap eligibility at something like active in the last two to four weeks, depending on your natural usage cadence. A weekly-use tool and a daily-use tool need different windows.

How many responses before you believe the number

Enough that a handful of people can't swing it. The same Learning Loop survey playbook puts the practical floor at 40 to 100 responses, and it's blunt about why: under 40, a few outliers can move your score by 10-plus points, which is the difference between "we have fit" and "we don't." Around 100 the number settles down.

Run the arithmetic and it's obvious. With 30 responses, one person flipping from "somewhat" to "very disappointed" moves the score by 3.3 points. Four people flipping moves it 13 points. You are not measuring your product at that sample size; you're measuring noise. Get to 100 and each response is worth one point, which is a resolution you can reason about.

If you want to slice the score by segment (plan tier, use case, acquisition channel), you need a lot more, because each slice is its own small sample. Budget 200-plus if segmentation is the goal. More on why segmentation matters in a second, because it's the part almost everyone skips.

Quick reference:

Response count What it's good for
Under 40 Basically a vibe. Don't quote a number.
40 to 100 A directional whole-product score
100 to 200 A score you can track over time and trust
200+ Enough to segment and still have signal per slice

The part everyone omits: mine the "very disappointed" open text

If you stop at the percentage, you've thrown away the most valuable thing the survey gives you. Right after the multiple choice, add one open field: "What's the main benefit you get from [product]?" Then read only the "very disappointed" answers first.

Those are your fans describing your value proposition in their own words. Not your positioning deck. Theirs. Patterns fall out fast: the same two or three phrases keep reappearing, usually not the ones your marketing leads with. That's the value you should double down on, because it's the value that's already creating people who can't live without the thing.

This is exactly how Rahul Vohra ran it at Superhuman. Their first score was a dismal 22%, nowhere near fit. Instead of despairing, his team segmented on the "very disappointed" group, learned who those people were and what they loved, then spent roughly half the roadmap deepening that and the other half converting fence-sitters. First Round Review's account of the engine tracks the climb from 22% to 58% over about a year. The survey wasn't a scoreboard. It was a roadmap-prioritization tool that happened to emit a number.

There's a companion move for the "somewhat disappointed" crowd: ask them what would make the product a must-have. Those are your fence-sitters, and their blockers are your growth backlog. But read fans first — you want to know what to protect before you know what to fix.

A worked example with actual numbers

Say you run the survey and 120 activated users answer.

  • Very disappointed: 44 (37%)
  • Somewhat disappointed: 58 (48%)
  • Not disappointed: 18 (15%)

37%. Below the line, but not by much, and 120 responses is enough that the number is real rather than jittery. A weaker team files this under "no fit, keep guessing." Here's the better read: you're close, and 58 somewhat-disappointed users is a fat middle to work with.

Now segment. Split the very-disappointed 44 by primary use case and you find that among users who came for, say, the bulk-import workflow, the very-disappointed rate is 61%, while among everyone else it's 24%. That's not one product with mediocre fit. That's strong fit with one audience diluted by users you're serving badly. The strategic question stops being "how do we raise 37%" and becomes "do we aim the product squarely at the import crowd, or invest to make the other use case actually land?" The aggregate number hid that entirely. The segment cut surfaced it.

Where this goes wrong

A few failure modes I've watched people walk into.

Treating 40% as a pass/fail gate. It's a heuristic from one person's dataset, not a law of physics. 38% with a steep upward trend and a rabid fan segment beats 42% that's been flat for three quarters. Track the trajectory, not just today's snapshot. Honestly, the industry doesn't fully agree here — some practitioners argue the 40% figure is over-cited and under-validated, and they have a point that it was never a controlled study. I still use it, because a rough-but-consistent yardstick you actually run beats a perfect one you never do. Just hold it loosely.

Surveying the wrong people to inflate the score. If you quietly restrict the survey to your happiest power users, you can manufacture any number you want. You'll also learn nothing. The score is only useful if the sample honestly represents your activated base.

Running it once. A single reading tells you where you are. It can't tell you whether you're climbing or sinking, which is the thing that actually predicts the future. Vohra ran it quarterly on purpose. Pick a cadence and hold it.

Reading the percentage and skipping the text. I'll keep saying it because it's the most common waste. The number tells you whether to celebrate. The open-text answers tell you what to build. If you only capture the first, you paid the full cost of running a survey for a fraction of the value.

If you're wiring PMF into a broader metrics setup, it slots in nicely next to a retention-driven north-star metric tree: the survey tells you whether you have a must-have, and the metric tree tells you whether that must-have-ness is showing up in behavior over time. One is stated preference, the other is revealed. When they agree, you can trust both. When they disagree, that gap is the most interesting thing on your desk this quarter.