Product Market Fit Framework: Four Models Evaluated on Accuracy

Product Market Fit Framework: Four Models Evaluated on Accuracy

Founders often treat one positive survey result, a rising activation rate, or a few large accounts as proof of fit. None of these proves durable demand. A product can score well with early users and still fail on retention, pricing, distribution, or unit economics.

The product market fit framework used to evaluate a company must match the company’s stage. Sean Ellis’s 40% survey rule can detect an early signal. Brian Balfour’s Four Fits can test whether growth mechanics align. Demand-first and supply-first models explain how a company reaches fit. First Round Capital’s maturity model tracks how fit develops. Retention cohorts determine whether the signal survives contact with time.

These frameworks do different jobs. Treating them as competing scorecards produces bad decisions.

A PMF survey measures perceived loss. Retention measures repeated value. Repeated value is the metric that survives.

The four product-market fit frameworks at a glance

The models below should not be ranked as if they answer the same question. They operate at different levels.

FrameworkPrimary questionBest useMain failure mode
Sean Ellis TestWould active users care if the product disappeared?Detecting early product demandSelf-selection, weak user definition, no retention proof
Brian Balfour’s Four FitsAre market, product, channel, and business model aligned?Evaluating scalable growthRequires more data and judgment
Demand-first vs. supply-firstDid the company discover demand or create a new supply pattern?Choosing a path to market entryCan become a narrative after the fact
First Round maturity modelHow far has PMF progressed?Tracking fit across company stagesStage labels can hide weak unit economics
Retention cohortsDo customers continue to receive value?Confirming sustainable fitSlow to mature and segment-sensitive

The first four help explain the system. Retention tells the company whether the system works.

The Sean Ellis Test: a useful early signal with a narrow scope

The Sean Ellis Test measures product-market fit through one user question. Active users are asked how they would feel if they could no longer use the product. If at least 40% select “very disappointed,” the product has a strong early PMF signal.

The rule is useful because it forces a company to measure perceived product value instead of collecting anecdotes from the loudest customers. It also produces a number that can be tracked across product changes, customer segments, and acquisition channels.

It does not produce a full PMF verdict.

The test becomes unreliable when the sample is poorly defined. A company can reach the threshold by surveying:

  • Users who already have high usage frequency.
  • Customers who received heavy onboarding support.
  • A narrow segment with unusually strong need.
  • Free users who do not represent the target buyer.
  • Employees, investors, or friendly early adopters.
  • Users acquired through a channel that cannot scale.

The sample needs an operational definition. “Active user” must mean something tied to value creation. For a collaboration product, that may require recurring team activity. For a B2B analytics platform, it may require repeated report consumption by the buying organization. A login is not a value event.

The survey also has a timing problem. Early users are not a random market sample. They often tolerate defects because the problem is urgent or because they helped shape the product. Their answers can confirm that the product matters to them without proving that the company can acquire similar users at an acceptable customer acquisition cost.

What the 40% threshold can and cannot tell you

The threshold is most useful as a leading indicator. It can answer:

  • Does the product create a strong reaction among a defined group?
  • Which segment reports the highest perceived loss?
  • Did a product change increase perceived value?
  • Is one use case stronger than the rest?

It cannot answer:

  • Whether users retain after the initial purchase.
  • Whether the product supports a viable pricing strategy.
  • Whether acquisition channels can scale.
  • Whether gross margin supports the business model.
  • Whether the person surveyed is the economic buyer.
  • Whether the product survives competitive substitution.

Slack is a useful example of the distinction. A 2015 test associated with Hiten Shah measured a 51% score among 731 users. That result exceeded the 40% benchmark. It showed strong user attachment. It did not, by itself, establish the full future economics of the company.

This is the core limitation of the Sean Ellis Test. It measures the cost of disappearance in the user’s mind. It does not measure the company’s ability to turn that value into durable revenue.

Sean Ellis Test versus the Superhuman PMF survey

The Sean Ellis Test is often discussed alongside the survey approach used by Superhuman. The logic is similar: identify users who would be highly disappointed by losing the product, then study what distinguishes them.

The distinction is operational. Superhuman’s approach became known for using the survey output as a segmentation and product-development tool. The company examined why users valued the product, which use cases correlated with strong responses, and where the product experience failed to meet expectations.

That is more useful than treating 40% as a pass-fail gate.

A company should split survey results by:

  • Customer segment.
  • Account size.
  • Role and buyer type.
  • Acquisition source.
  • Product frequency.
  • Time since activation.
  • Paid versus free status.
  • Core use case.

A blended score can hide a business with strong fit in one segment and no fit elsewhere. A 45% overall score may be less useful than a 32% score concentrated in a segment with high retention, expansion revenue, and low support cost.

The correct reading is simple: the survey identifies where to investigate. It does not close the investigation.

Brian Balfour’s Four Fits: the strongest model for scaling mechanics

Brian Balfour’s Four Fits Framework expands PMF beyond the product-market boundary. It argues that a company cannot reach large-scale growth through product value alone. Four forms of alignment must develop:

1. Market-Product Fit — the product solves a meaningful problem for a defined market.

2. Product-Channel Fit — the product can be distributed through a channel that matches how users discover and adopt it.

3. Channel-Model Fit — the channel economics support the business model.

4. Model-Market Fit — pricing, monetization, and market structure align.

This framework addresses the failure mode that the 40% survey rule ignores. A product can create real value and still fail to scale because acquisition costs rise faster than lifetime value, the sales cycle exceeds cash capacity, or the product requires onboarding that the chosen channel cannot support.

Market-Product Fit

This is the layer most founders call PMF. The market has a problem. The product addresses it. Users return or pay.

Evidence includes:

  • Retention cohorts that flatten above zero.
  • Repeated use of a core workflow.
  • Organic referrals or direct demand.
  • Expansion within accounts.
  • Willingness to pay without founder-led persuasion.
  • A defined segment with consistent use cases.

A high survey score supports this layer. It does not settle it.

Product-Channel Fit

A product must match the behavior of its acquisition channel. Self-serve software requires a product that communicates value without a sales engineer. Enterprise software can tolerate a sales process but must justify implementation cost. A consumer product may rely on sharing, habit, or network effects.

Channel mismatch creates false negatives and false positives.

A technically strong product can fail through paid acquisition if the activation event occurs after a long setup process. A product can also produce strong organic growth in a small community that does not represent the broader market.

The useful question is not whether a channel produces leads. It is whether the product reaches its value event through that channel at a repeatable cost.

Channel-Model Fit

The acquisition channel must work with the revenue model.

A sales-led motion can support high contract values and implementation costs. It usually cannot support low annual pricing with long procurement cycles. A self-serve channel can support lower contract values when activation is fast and support load remains low. It breaks when every customer requires custom work.

This is where CAC payback, gross margin, sales capacity, and burn multiple enter the analysis.

A product with a 40% survey score can still be a poor venture business if:

  • CAC exceeds the gross profit available from the customer.
  • Payback depends on renewal that has not been observed.
  • Sales productivity falls as the founder stops closing deals.
  • Support costs increase with every new account.
  • Expansion revenue is assumed rather than measured.

Model-Market Fit

The business model must fit the market’s purchasing behavior.

Some markets buy seats. Others buy usage, transactions, outcomes, or access to regulated infrastructure. Pricing by the wrong unit creates friction. It can also distort retention analysis. A customer may remain active while reducing consumption, or appear retained while the account produces no margin.

Model-market fit requires alignment between:

  • Value delivered.
  • Unit charged.
  • Budget owner.
  • Purchase frequency.
  • Contract duration.
  • Expansion path.
  • Cost to serve.

The Four Fits model is more demanding than a survey because it forces the company to connect demand with distribution and economics. It is not a faster diagnostic. It is a better scaling diagnostic.

Product-market fit is local. Scalable fit is a chain. One broken link caps growth.

Demand-first versus supply-first: two paths to market entry

Casey Winters frames PMF strategy through two paths. The demand-first path starts with an existing customer problem. The supply-first path starts with a product vision or a new capability that changes what users can do.

The distinction matters because the evidence required from each path differs.

Demand-first: start with customer pain

The demand-first model resembles the Eric Ries approach. The company observes a problem, builds a narrow solution, collects usage data, and iterates.

Its advantages:

  • The initial problem is visible.
  • Customer interviews can identify the workflow.
  • Early users provide direct feedback.
  • The company can test willingness to pay before building a broad platform.
  • Product scope remains constrained.

Its risks:

  • The company builds a feature instead of a business.
  • Early customers define a custom product that cannot generalize.
  • The market appears large because many people report the problem, but few pay to solve it.
  • The product becomes a services operation.
  • The company optimizes for requests instead of retention.

Demand-first execution requires discipline around the core use case. Each customer request should be tested against retention, expansion, and repeatability. A request from one high-value account is not automatically a roadmap priority.

The relevant evidence is behavioral. Users must adopt the workflow without constant founder intervention. The business must show that the same value proposition works across a defined segment.

Supply-first: create a new usage pattern

The supply-first model is associated with the Keith Rabois view. The company starts with a product insight or a new form of supply, then develops demand around it.

This path appears in products that enable behavior customers did not previously request because the category did not exist in a familiar form. The user cannot always describe the need in advance. The product must demonstrate the use case.

Its advantages:

  • The company can create a new market category.
  • Product differentiation may be stronger.
  • The product can avoid direct feature comparison.
  • New supply can unlock demand that legacy providers do not serve.

Its risks are severe:

  • Education costs remain high.
  • The initial market may be too narrow.
  • Usage can reflect curiosity rather than durable value.
  • Distribution may require a channel that does not exist.
  • Investors and founders can confuse novelty with PMF.

Supply-first companies need stronger behavioral proof. A survey can capture enthusiasm. It cannot distinguish durable adoption from early interest. Retention, repeat usage, referral behavior, and willingness to pay carry more weight.

Choosing the correct path

The two models are not mutually exclusive. A company can begin with an observed problem and introduce a new product behavior. It can also launch a new technology and discover a customer pain that becomes the core market.

The useful distinction is the source of the initial hypothesis:

Entry pathInitial evidenceMain testCommon error
Demand-firstCustomer pain and existing workflowCan the product improve a repeated job?Building custom features for each buyer
Supply-firstNew capability or product behaviorCan the company create repeated demand?Treating novelty as proof of value

Demand-first companies usually need to prove generalization. Supply-first companies usually need to prove persistence.

First Round Capital’s maturity model: PMF is not binary

First Round Capital describes PMF through four maturity levels:

1. Nascent PMF

2. Developing PMF

3. Strong PMF

4. Extreme PMF

This model solves a common management error: treating PMF as a switch that is either on or off.

A product can have real demand in a small segment and still lack the conditions for broad scaling. It can have strong retention but no efficient channel. It can have a repeatable acquisition channel but weak monetization.

Nascent PMF

Nascent PMF means the company sees evidence of demand but cannot yet explain it with precision.

Typical signals:

  • A small group of users returns.
  • Some customers report high value.
  • The core use case is still changing.
  • Retention data is incomplete.
  • Acquisition depends on founders or personal networks.
  • Revenue is too concentrated to establish a stable pattern.

The operating priority is discovery. The company should reduce uncertainty about the user, problem, and activation event. Scaling spend at this stage tends to convert uncertainty into burn.

Developing PMF

Developing PMF means the company has a repeatable use case in a defined segment. The product retains some users, but the model still has weak points.

Typical problems:

  • Retention varies by cohort.
  • One channel works while others fail.
  • Pricing is not settled.
  • Support or onboarding remains manual.
  • The buyer and user are not always the same.
  • Expansion revenue is inconsistent.

This is the stage where the Four Fits framework becomes useful. Product value exists, but the company must determine whether distribution and monetization can carry growth.

Strong PMF

Strong PMF means the company has evidence across product usage, retention, acquisition, and revenue. Growth is not dependent on one founder, one account, or one temporary channel.

The company can usually answer:

  • Which segment retains?
  • Which event predicts renewal?
  • Which channel acquires that segment?
  • What is the payback profile?
  • Which pricing unit matches delivered value?
  • What breaks when volume increases?

Strong PMF does not mean the company can spend without control. It means the growth system has enough repeatability to justify investment.

Extreme PMF

Extreme PMF is rare. The research framework identifies it as a condition reached by roughly 5% to 10% of startups. The figure should be read as a classification signal, not as an audited market statistic.

Extreme PMF usually combines:

  • High retention.
  • Strong organic demand.
  • Efficient or improving acquisition.
  • Low friction in adoption.
  • Expansion behavior.
  • Clear pricing power.
  • A market that can absorb much more supply.

At this stage, the constraint shifts. The question is no longer whether customers want the product. The question is whether the company can deliver, hire, finance, and govern growth without destroying the economics.

The maturity model is valuable because it prevents premature scaling. It also prevents the opposite mistake: rejecting a product because it has not reached extreme fit while it is still developing a strong segment.

Retention cohorts: the final test

Cohort retention is the strongest quantitative evidence of sustainable product-market fit. A retention curve that flattens above zero shows that a segment continues to receive value over time.

The curve matters more than the first-month peak. Many products generate an initial burst of use. The business case depends on what remains after novelty, onboarding, and promotional activity end.

A cohort analysis should separate:

  • Signup cohorts.
  • Paid conversion cohorts.
  • Customer segments.
  • Acquisition channels.
  • Contract sizes.
  • Use cases.
  • Geography.
  • Product plans.
  • Buyer roles.

A blended retention curve can conceal a failing segment behind a successful one. It can also hide a channel problem. Customers acquired through a partner may retain while paid-search customers churn. The aggregate number will not explain why.

What a healthy retention curve shows

A useful cohort curve does not need to retain every user. The relevant standard depends on the business model.

A consumer subscription product, a B2B platform, and a transaction-based service have different retention mechanics. But sustainable fit requires a remaining base that continues to generate value or revenue. A curve that falls to zero indicates that initial adoption did not become a durable behavior.

For B2B products, logo retention alone can mislead. An account may renew while seats decline, usage falls, or support costs rise. Revenue retention and usage retention should be reviewed together. Net revenue retention can improve while the product loses smaller accounts. That may be acceptable if the target segment is changing, but the tradeoff must be explicit.

For usage-based businesses, account retention can hide consumption decline. The company needs to track the unit that drives revenue.

Retention and the 40% rule together

The survey and cohort approaches answer different questions:

MetricWhat it measuresSignal strengthRequired follow-up
“Very disappointed” responsePerceived product lossEarly demand indicatorSegment users and compare with retention
Activation rateInitial value realizationProduct onboarding signalLink activation to later retention
Repeat usageContinued engagementBehavioral fit signalIdentify the workflow that creates recurrence
Logo retentionContinued customer relationshipBasic B2B durabilityAdd revenue, usage, and seat analysis
Revenue retentionPreserved or expanded revenueCommercial durabilityCheck margin and account concentration
Cohort curvePersistence over timeStrongest PMF evidenceSegment by channel, use case, and customer type

A company that passes the 40% rule but fails cohort retention has a perception signal without durable fit. A company with strong retention but a weak survey score may have a product that creates value without generating emotional attachment. That can still be a viable business, especially in infrastructure and workflow categories.

The metrics should be read together. No single number carries the entire decision.

How to evaluate a product market fit framework in practice

A framework is accurate when it reduces decision error. The practical test is whether it helps a team decide what to do next without hiding uncertainty.

The four models perform different functions:

1. Use the Sean Ellis Test to locate strong user value.

Survey active users. Define active by value-producing behavior. Segment the results. Treat the 40% threshold as a leading indicator.

2. Use the Four Fits to test scale mechanics.

Map the product to its acquisition channel, the channel to the business model, and the model to the market. Identify the weakest fit.

3. Use demand-first or supply-first to evaluate the company’s entry logic.

Demand-first products need proof that the solution generalizes. Supply-first products need proof that new behavior becomes repeated demand.

4. Use the maturity model to classify the current stage.

Do not call nascent evidence strong PMF. Do not demand extreme PMF before investing in a proven segment.

5. Use retention cohorts to approve or reject scaling.

A retention curve that flattens above zero carries more weight than a polished survey result, a large pipeline, or a founder’s conviction.

This sequence also improves capital allocation. Early survey data may justify more discovery work. Cohort retention and channel economics may justify a larger growth budget. A high burn multiple with weak retention does not become rational because the survey score is above 40%.

The right PMF framework is not the one with the cleanest score. It is the one that exposes the next failure mode.

The verdict

The Sean Ellis Test is the best lightweight instrument for detecting early user attachment. It is fast, cheap, and easy to repeat. It is not a survival test.

Brian Balfour’s Four Fits is the strongest framework for evaluating whether a product can scale beyond a narrow pocket of demand. It connects product value to distribution and economics.

The demand-first versus supply-first model is useful for understanding how the company must prove its thesis. It is a strategic lens, not a measurement system.

First Round Capital’s maturity model is the best structure for avoiding binary PMF thinking. It shows that fit develops through stages and that each stage has a different operating requirement.

Retention cohorts are the final arbiter. A product that retains customers, supports acceptable economics, and reaches them through a repeatable channel has viable fit. A product that produces enthusiasm without persistence does not.

The binary verdict is straightforward:

Use surveys to find the signal. Use Four Fits to test scale. Use retention to decide whether the business is real.

FAQ

What does the Sean Ellis 40% rule measure?
It measures how users would feel if they could no longer use the product. If at least 40% of active users select “very disappointed,” the result indicates a strong early product-market fit signal, not a complete verdict.
Why is the Sean Ellis Test not enough to prove product-market fit?
The test measures perceived product value among surveyed users, but it does not show whether customers retain, pay sustainably, can be acquired through scalable channels, or support viable unit economics.
What are Brian Balfour’s Four Fits?
The Four Fits are Market-Product Fit, Product-Channel Fit, Channel-Model Fit, and Model-Market Fit. Together, they assess whether product value connects with distribution and business economics.
What is the difference between demand-first and supply-first product strategies?
Demand-first starts with an existing customer problem and develops a solution around a repeated workflow. Supply-first starts with a new capability or product behavior and must create repeated demand around it.
Why are retention cohorts important for product-market fit?
Retention cohorts show whether customers continue to receive value after initial adoption, onboarding, or promotional activity ends. A retention curve that flattens above zero is stronger evidence of sustainable fit than an initial survey result.
How should companies use product-market fit frameworks together?
Use the Sean Ellis Test to locate strong user value, Four Fits to test scaling mechanics, the demand-first or supply-first lens to evaluate market-entry logic, the maturity model to classify the current stage, and retention cohorts to decide whether scaling is justified.