You’ve read that you’ll know product market fit when you feel it, that it’s the pull rather than the push, servers straining, customers buying faster than you can hire. That’s true, and it’s useless on a Tuesday morning when your board asks whether you have it and your honest answer is “i think so?”
This article is about measuring product-market fit, not describing it: the survey, the retention curve, a signals table, four Indian examples, what changes for AI products, and what to do when the honest answer is no. PMF isn’t a switch that flips either, it’s a spectrum, specific to a segment, and it can be lost. Worth saying upfront, since most pages implying the opposite.
What Product-Market Fit Actually Means
Product-market fit is the point at which a product satisfies a strong existing demand in a specific market segment well enough that customers keep using it, keep paying for it, and tell others about it, without the company pushing them at every step.
Two things in that sentence are measurable, and that’s the whole article: a market, and a degree of satisfaction. Quick distinction worth having: problem-solution fit means you’ve confirmed the problem is real and your approach addresses it. Product-market fit means the market’s behaviour proves it at scale.
The lineage, correctly attributed for once: Don Valentine, Sequoia’s founder, developed the underlying market-first thinking. Andy Rachleff, Benchmark’s co-founder, named the concept. Marc Andreessen popularised it in the mid-2000s, crediting Rachleff. The concept was built by investors, for investors, to decide where to put money, which is exactly why the canonical writing is impressionistic rather than operational. Founders and PMs need the same idea from the inside, as an instrument, not a vibe. That’s this article’s whole thesis. (See Wikipedia’s sourced attribution chain if you want the full history.)
What it isn’t, quickly: not a funding round (investors price on narrative too), not a launch or a press cycle (those measure attention), not raw signup growth (growth you bought isn’t evidence of anything), and not a single number crossing a threshold, including any number in this article.
Andreessen’s description, servers straining, hiring nonstop, is real, but it only describes the extreme. Most founders live in the ambiguous middle where the signal is real but quiet, and “you’ll know” is useless there. It’s also unfalsifiable, which makes it unmanageable, you can’t set a goal against a feeling. Rahul Vohra’s contribution at Superhuman was turning fit into a number he could put in an OKR. Judging fit well is also, not incidentally, a core part of what a product manager’s job actually involves day to day.
How to Measure Product-Market Fit
No single metric proves fit. You’re triangulating a stated-preference signal (the survey), a revealed-preference signal (retention), and a social signal (organic growth). When they disagree, trust retention, it’s behaviour, not opinion.
The Sean Ellis Test, and What 40% Really Means
The question: “how would you feel if you could no longer use [product]?” Three options, very disappointed, somewhat disappointed, not disappointed. The score is the percentage answering very disappointed.
The benchmark is 40%, and here’s the attribution most pages skip: Sean Ellis, an early growth leader at Dropbox, LogMeIn and Eventbrite, arrived at it surveying roughly 100 startups, below 40% they struggled to find sustainable growth, above it they tended to have traction. It’s a heuristic from a small, self-selected sample, not a statistical law. 39% isn’t failure and 41% isn’t success. The trend across cohorts matters more than the exact number.
Survey people who’ve genuinely used the product recently, at least twice in the last two weeks, not everyone who ever signed up, that measures onboarding, not fit. Aim for 100+ total responses and 40+ within any segment you want to draw a conclusion about. Two follow-ups matter: who’d benefit most, and how could we improve it, the first gives you the segment in users’ own words, the second gives you the roadmap.
Superhuman’s first survey scored 22%. Rather than treating that as a verdict, Rahul Vohra’s team segmented respondents by persona, the score jumped to 33% immediately from segmentation alone, then to 58% over three quarters of splitting the roadmap 50/50 between what fans already loved and what held fence-sitters back. Below ~25%, something structural is wrong. Between 25 and 40%, look for a segment already scoring high and consider narrowing. Above 40%, your job shifts from finding fit to not losing it while you scale.
Retention Curves That Flatten: The Strongest Signal You Have
Take everyone who signed up in a given month, plot what percentage are still active one month later, two months later. You’re not looking at the curve’s height. You’re looking at its shape. A curve that keeps falling means every user eventually leaves, a leaky bucket. A curve that flattens means a stable share made your product a habit, that plateau is the clearest evidence of fit that exists, because it’s behaviour.
| Months after signup | Product A (active) | Product B (active) |
| Month 1 | 42% | 45% |
| Month 3 | 27% | 19% |
| Month 6 | 24% | 6% |
| Month 8 | 23% | 2% |
Product B looks better at month 1. Founders judging on a 30-day funnel would pick B. By month 6 the gap is enormous, A has flattened around 23–24%, B is heading to zero. You cannot assess product-market fit from a 30-day number, arguably the most useful sentence in this article.
Run both for a year at 1,000 signups a month. Product A’s active base compounds to roughly 3,880 users by month 12 and keeps climbing by about 230 users every month indefinitely, each new cohort deposits a permanent ~23%. Product B plateaus around 2,270 and stops, new cohorts arrive at exactly the rate old ones evaporate. B’s founder can only grow by buying more users; pause spend and the base falls. A’s founder owns a compounding asset. Same category, same acquisition spend, opposite businesses. Double B’s acquisition and the base roughly doubles, then stops again, at twice the cost. Growth that needs permanent spending isn’t evidence of fit.
Rough six-month retention reference points from Lenny Rachitsky and Casey Winters (2020, worth the date caveat): ~25–45% for consumer social, ~40–70% for consumer SaaS, ~75–90% for enterprise SaaS. Treat these as category medians, not thresholds, and judge your curve against the real-world frequency of the problem, not a generic number, a tax app retaining 12% monthly may have excellent fit; a messaging app at 60% may not.
Organic Growth, and Why NPS Won’t Answer This
If users bring you more users unpaid, the market is doing your distribution, the social confirmation of what the retention curve already shows. Measure the share of signups arriving unpaid, and run the uncomfortable test: pause paid acquisition for two weeks and see what happens. Most founders never run it. It’s extremely informative, and it’s a lagging indicator, weak in low-frequency B2B categories where the equivalent signals are shortening sales cycles and buyers arriving already informed.
Net Promoter Score measures satisfaction and willingness to recommend. The Sean Ellis question asks whether losing you would actually hurt, a much harder question to answer politely. NPS is inflated by response bias and skewed by timing. Useful for tracking satisfaction over time, not for deciding whether you have fit.
Segments Hide What Averages Erase
A blended 22% score can contain one segment at 55% and three at 8%. The blended number says give up. The segmented number tells you exactly who to build for, that’s Superhuman’s entire method. Segment by role, company size, channel, use case; you need 40+ responses per segment for the number to mean anything, and it’s easy to keep slicing until some sub-group looks good, a high score in a segment too small to build a business on is a curiosity, not a strategy. Sizing that promising segment before committing is its own discipline, worth handing off to elsewhere rather than repeating here.
Leading indicators move before fit shows up in revenue, activation rate, week-2 return rate, depth of use. Lagging indicators confirm what already happened, revenue growth, blended churn, press coverage. Manage on the leading ones, report on the lagging ones, which is standard practice in analytical business roles more broadly, not a startup invention.
The Product-Market Fit Signals Table
No single row is decisive. The diagnosis comes from how many agree, and rows that disagree tell you where to look next.
| Signal | You have it | You don’t | How to measure |
| Survey score | 40%+, stable or rising | Below 25%, or high only in a tiny segment | In-product survey, active users, 100+ responses |
| Retention curve | Flattens by month 3–6 and holds | Keeps declining toward zero | Cohort table by signup month, 6+ months |
| Organic growth | Large, rising unpaid share | Growth stops within days of pausing ads | Channel attribution; a paid-pause test |
| Depth of use | Adopts beyond the entry feature | Shallow, single-feature use only | Feature-adoption breadth |
| Behaviour in an outage | Users complain loudly and wait | Silence, then quiet churn a month later | Support volume vs post-incident churn |
What Changes in the AI Era
The definition hasn’t changed. Where the difficulty sits has. For twenty years the hard part of testing an idea was building it. That’s no longer reliably true, and three consequences follow.
In March 2025, YC managing partner Jared Friedman told TechCrunch that for roughly 25% of the Winter 2025 batch, over 95% of the codebase was AI-generated, excluding library imports, one accelerator batch, a direction indicator rather than a global statistic. When four people can ship a credible product in a fortnight, speed stops being a moat, because everyone has it. The scarce inputs become distribution and judgement, both always mattered, they were just partly hidden behind how hard building used to be. Practical result: the binding constraint shifts from build capacity to learning capacity. Six half-measured experiments teach you less than two properly instrumented ones, the measurement discipline above is more valuable now, not less.
AI products have a specific failure mode conventional software doesn’t: the demo genuinely delights and production genuinely doesn’t. A demo curated on clean inputs can fall apart on the long tail of real ones, ambiguous requests, messy documents, edge cases nobody anticipated. Every early signal gives a false positive, trial signups are high because the demo compels, week-1 engagement is high because people are curious, and month 3 arrives with the users who tried real workloads already gone. Stop asking whether it works on the average input; ask whether it holds at the hardest 5% of real ones, that’s what users actually judge you on. The widely-cited MIT NANDA finding that roughly 95% of enterprise GenAI pilots showed no measurable P&L impact is worth knowing, though its methodology has been publicly questioned, treat it as a directional signal, not a settled fact. The underlying point holds regardless: pilots convert to production far less than demo enthusiasm predicts.
A product that’s a thin interface over a general model captures value only while that interface beats the model’s own. A useful test: if the model underneath gets 30% better next quarter, does your product become more valuable or less necessary? If less necessary, you’re renting your fit from someone else’s roadmap. What sits above the model: the workflow around the output, proprietary data, distribution, and evaluation infrastructure that proves reliability to a buyer. This is exactly why retention matters more than growth for AI products specifically, growth can come from novelty, and novelty has a shelf life measured in weeks.
Practically: measure retention from month 3 onward, not month 1, a16z’s research on “AI tourists” argues early AI cohorts are inflated by curious users who churn out fast. Watch churn by price tier too, ChartMogul’s analysis of AI-native companies found gross revenue retention swinging heavily with price, roughly 70% above $250/month versus ~23% below $50, small sample, real direction. Track task success rate and human-correction rate as close-to-real-time fit signals, and track gross margin per retained user separately, a cohort that retains beautifully but loses money on inference is a different problem than one that churns, and it’s invisible on a standard retention chart.
Four Indian Products and the Signal That Told Them They Had It
Worth a caveat before these: these are companies that worked, we’re reading their signals backwards, and plenty of companies had similar early signals and still failed. The value is the shape of the signal, not a formula.
Meesho started as Fashnear, hyperlocal fashion delivery. Within six to nine months, more than half its users were a segment nobody designed for, micro-entrepreneurs reselling through WhatsApp and Facebook. Per Elevation Capital’s account of its August 2017 investment, early traction was roughly 4,400 registered entrepreneurs and 100%+ month-on-month growth. They rebuilt the entire product around the resellers who showed up uninvited. The lesson: the most valuable thing in your data is often the segment you didn’t plan for.
Zerodha, discount broking with flat per-order pricing in a market charging by percentage, grew for years almost entirely without advertising, through referrals and its own educational content (worth noting Groww overtook it on active-client count in late 2023, so check current standing before claiming “largest”). The insight sat in the pricing model, not the interface, flat pricing changed who could afford to trade actively, which changed the market instead of just competing in it.
Zomato began as Foodiebay, scanned restaurant menus shared internally at a consulting firm because colleagues were tired of paper menus. They kept coming back repeatedly, unasked, a retention-shaped signal inside an audience of a few hundred, long before there was a market. The earliest real evidence of fit is often tiny and behavioural, repeat use, not enthusiasm, and the difference from over-reading friends’ feedback is that this was unprompted repeated use, not a compliment.
Razorpay, a developer-first payments API where getting a gateway traditionally meant paperwork and a sales process, grew through developers integrating without a salesperson involved, and recommending it to other developers. In a B2B category, self-serve adoption by a technical buyer is the enterprise dialect of word of mouth, measurable in time-to-first-successful-API-call and the share of revenue from accounts nobody sold to.
What to Do When You Don’t Have Product-Market Fit
Most companies at any given moment don’t have it, including many that eventually will. A measurement that says not yet is a useful result, not a verdict. The actual failure is not knowing.
Diagnose before you pivot, changing everything at once destroys the information you’d need to learn from the change. Segment the survey and retention data first, is the signal weak everywhere or strong somewhere small? Interview your ten most-retained users, not your loudest, find out what job they’re really using it for. Interview ten who churned after genuinely trying it, they know exactly what’s wrong. Then separate the failure point: wrong segment, unimportant problem, wrong mechanism, or a distribution problem masquerading as a fit problem, strong retention with weak growth is almost always distribution; weak retention with strong growth is almost always fit.
| Pivot type | What changes | Evidence pointing here |
| Customer segment | Who you build for | One unplanned segment retains far better than your target one |
| Problem | Which job you solve | Users like you, but the problem isn’t painful enough to prioritise |
| Solution | The mechanism | Everyone agrees the problem’s real; nobody sticks with your approach |
| Business model | How you charge | People love using it; the money doesn’t work, or price blocks adoption |
Change one variable and time-box it, a full retention cycle, 8–12 weeks minimum, longer for low-frequency products, before judging, the signal doesn’t appear in 30 days, as the worked example above showed. Four questions worth an honest afternoon: is any segment above 25% on the survey, that’s your cheapest pivot, narrow into it. Has any cohort’s curve ever flattened, anywhere? Can you find ten people who’d be genuinely very disappointed to lose it, if not, there’s no niche to expand from yet. What’s the cheapest experiment that would actually change your mind, if it costs six months and can’t be falsified, it’s a bet, not an experiment.
And, honestly, when to stop: no segment crossing ~25% after three genuine rounds of change, no cohort ever flattening across four-plus cohorts, every growth increment needing proportional spend with zero compounding, under six months of runway with nothing moving. Stopping is a resource-allocation decision, not a character judgement, your time is the scarce asset, and the same team with better information is often the strongest thing to redeploy.
The Honest Limitations of Product-Market Fit
It isn’t binary, it’s a spectrum. Superhuman at 22%, 33% and 58% was the same company measuring the same thing at different intensities, the useful question is never do we have it but how much, with whom, and is it rising.
It’s segment-specific, always, you never have fit with “a market,” only with a describable group sharing a job to be done, which is why expansion is genuinely risky, fit doesn’t automatically transfer to an adjacent segment.
It can be lost, usually quietly, not dramatically, a competitor arrives, expectations shift, or in 2026 specifically, a model release can eliminate a product’s reason to exist between one quarter and the next. The survey and retention curve are a standing instrument, not a one-time exam.
And every measurement here can be gamed, often by accident. Survey only your most engaged users and the score rises. Define “active” loosely and retention flattens on paper. Judge on 30 days and a decaying curve looks fine. Product-market fit isn’t a destination you arrive at, it’s a reading you take, on a specific segment, at a specific moment, and the discipline is taking it often enough to notice when it changes.
FAQs
What is product-market fit in simple terms?
When a product satisfies real, strong demand in a specific group well enough that they keep using it, paying for it, and telling others, without the company pushing at every step. A spectrum, not a switch.
How do you measure product-market fit?
Triangulate a survey (how disappointed would users be to lose it), a cohort retention curve (does it flatten), and organic growth share. When they disagree, trust retention, it’s behaviour, not opinion.
What is the 40% rule for product-market fit?
Sean Ellis found startups where 40%+ of users said they’d be “very disappointed” without the product tended toward sustainable growth. A heuristic from roughly 100 startups, a useful benchmark, not a statistical law.
What are the signs you have product-market fit?
A retention curve that flattens into a plateau, growth that survives pausing paid acquisition, users adopting beyond the entry feature, and loud complaints when the product breaks. Any single sign can mislead; look for several agreeing.
Is NPS a good measure of product-market fit?
Not alone. NPS measures satisfaction and willingness to recommend; PMF is about need. Users can rate you highly and still not miss you if you disappeared. Track NPS for satisfaction trends, not for a fit verdict.
What should you do if you don’t have product-market fit?
Segment your survey and retention data, interview your most-retained and recently-churned users, and identify whether the problem is segment, problem, solution, or distribution. Change one variable and give it a full retention cycle before judging.
Can a company lose product-market fit?
Yes. Markets shift, competitors with better distribution arrive, and for AI products specifically, a single model release can remove a product’s reason to exist. Treat measurement as ongoing, not a one-time exam.
Is product-market fit different for AI products?
The definition’s the same; measurement is harder. AI demos generate enthusiasm that inflates early signals, so trial and week-one numbers produce false positives. Measure retention from month three, and track task success and correction rates.
Wrapping Up
The essays are right that fit feels like pull. But you can measure the pull, with a three-question survey, a cohort retention table, and an honest look at what happens when you stop paying for growth. This week: run the survey on genuinely active users, plot six months of cohort retention, and segment both. Cheap building has made the measuring, not the making, the hard part.
Reading product signals honestly, segmentation, retention analysis, pricing, knowing when the data says change direction, is business judgement, and it’s learnable. Scaler’s online PGP in Business & AI is built around exactly this kind of decision-making for working professionals.
