Most product teams can tell you, in detail, what they built last quarter. Fewer can tell you, with any confidence, what changed for a customer because of it. That gap, between activity and impact, between output and outcome, is the subject of this guide. It lays out the Outcome-Driven Product Design Framework in full: twelve stages that move from identifying a real customer and a real problem, through building the smallest useful thing that can test a hypothesis, to measuring what actually happened, deciding what to do next, and eventually retiring the product responsibly when its value has run its course. This is the reference article. Other pieces on this site go deeper into individual stages; this one explains how they fit together and why the order matters.
The order matters because most product failures are not failures of execution. Teams that build carefully and ship reliably still produce things nobody needed, because the decision about what to build was made before anyone rigorously established who the customer was, what they were struggling with, and what evidence would tell you the struggle had actually eased. Outcome driven product design puts that decision first, in writing, before a single line of code or a single design mock is produced.
Output Is Not the Same Thing as Outcome
The distinction underneath this entire framework is simple to state and easy to lose track of under deadline pressure. Output is what a team creates: a feature, an app, a dashboard, a workflow, a report, an AI assistant. Outcome is what changes for the customer as a result: time saved, a problem resolved, confidence increased, a behavior changed, a risk reduced, a decision made better, an error reduced, revenue increased.
Output is what we make. Outcome is what becomes better because we made it.
Joshua Seiden made a similar point the center of his book “Outcomes Over Output,” defining an outcome as a measurable change in customer behavior that drives business results. Melissa Perri gave the failure mode a name in “Escaping the Build Trap”: organizations that measure themselves by what they ship rather than by what shipping accomplished. Marty Cagan and the SVPG community describe the organizational version of this problem as the difference between a “feature team,” which takes requirements and builds them, and an empowered “product team,” which owns a problem and is accountable for whether it actually gets solved. None of these are new observations. What this framework adds is a concrete, repeatable sequence for acting on the distinction rather than just nodding along with it.
Why This Matters More Now, Not Less
For most of software’s history, the cost of building the wrong thing was partially offset by how slow building anything was. A misguided feature took months to fully materialize, and that slowness created natural checkpoints where a team might notice, before too much was sunk, that the premise was shaky. AI-assisted development has removed much of that friction. GitHub’s own usage data and McKinsey’s research on generative AI both point to real, measurable gains in how fast software and prototypes can be produced. A working application that once took a small team a quarter can now be scaffolded by one or two people in days.
None of that speed makes a team better at knowing what to build. One is execution, the other is judgment, and a coding assistant has an opinion about neither the customer nor the problem. GitClear’s 2025 research, examining code commits from 2020 through 2024 as AI coding assistants became standard tooling, found rising code churn and growing code duplication over that period, consistent with work being produced faster than it is being understood. CB Insights’ recurring analysis of startup post-mortems has long found “no market need” as the single most cited reason founders give for failure, cited in roughly 42 percent of cases reviewed, well before generative AI existed. Slow execution was rarely the binding constraint on product success. Understanding the customer well enough to build the right thing was, and still is.
Velocity without direction creates waste faster. AI did not invent that risk. It just removed the friction that used to slow it down enough to notice.
That is the argument for a disciplined framework now more than ever: not because building is bad, but because the strategic question has shifted from “can we build it?” to “should we build it, for whom, and why?” Outcome driven product design is a structured way to keep asking that second question before the first one gets answered by default.
A Running Example
To keep the twelve stages concrete rather than abstract, this guide follows one illustrative, hypothetical case throughout: a small team at a fictional company, Northbridge Health, building a medication reminder feature for an existing patient portal app. Northbridge is not a real company and the example is constructed purely to demonstrate how the stages connect.
The 12 Stages
1. Customer Segments
Every product decision implicitly favors somebody, so it is worth making that choice explicit rather than pretending a product serves everyone equally. The framework distinguishes a core customer, roughly 80 percent of design attention as a guideline rather than a strict rule, a secondary customer who receives real but lesser consideration, and a broader population or indirect beneficiary who is affected without being a direct user. A product that tries to optimize equally for everybody often becomes weak for everybody. For Northbridge, the core customer is defined as patients over 60 managing two or more chronic prescriptions, the secondary customer is family caregivers who sometimes manage medications on a patient’s behalf, and the broader population includes the prescribing physicians who benefit indirectly from better adherence data.
2. Customer Needs, Pains & Outcomes
Not every complaint deserves a response, so the framework asks teams to assess Problem Strength before committing: frequency, how often the problem occurs; severity, how much it hurts when it does; urgency, how much it demands immediate attention; cost of inaction, what happens if nothing changes; and willingness to change, whether the customer will actually adopt something different. Alongside that, the Outcome Chain traces cause and effect from output to impact: Immediate Outcome, the first thing that changes; Behavior Change, what the customer starts or stops doing; User Success, the problem genuinely resolved from the customer’s point of view; and Long-Term Impact, the durable effect over time. For Northbridge, the immediate outcome might be “patient sees the reminder,” but the behavior change that actually matters is “patient takes the medication on time,” and user success is “adherence improves enough that the physician notices at the next visit.”
3. Assumptions & Evidence
Teams routinely write assumptions into product documents in a tone that makes them sound like facts, and that habit is worth breaking deliberately. The framework sorts what a team believes into four categories: Known Facts, verified with evidence; Assumptions, treated as true but not yet checked; Hypotheses, framed explicitly for testing; and Unknowns, gaps nobody has an answer for yet. An assumption written in a product document does not become a fact. Each key assumption needs an evidence source, a confidence level, a validation action, and a stated consequence if it turns out to be wrong. Northbridge’s team assumes patients will engage with push notifications; the validation action is a two-week pilot with fifteen patients, and the consequence if wrong is a pivot to a simpler channel such as a phone call or a printed schedule.
4. Minimum Viable Output
A Minimum Viable Output is not a smaller version of the finished product. It is the smallest, most focused output that can create, test, or meaningfully advance the priority outcome, which is a narrower and more disciplined idea than the popular sense of “MVP” as a stripped-down version of the eventual app. The lineage matters here and is worth crediting accurately: Eric Ries’s “The Lean Startup” and Steve Blank’s Customer Development work established the build-measure-learn discipline this stage draws on, but a Minimum Viable Output is specifically scoped to one outcome hypothesis rather than to an entire product vision. Teams sort candidate work into four buckets: Must Have Now, Can Wait, Will Not Build Now, and Dependencies. For Northbridge, the Minimum Viable Output is not a full medication management platform with refill ordering, interaction warnings, and caregiver dashboards. It is a single daily text message reminder tied to one prescription, sent at a time the patient chose during onboarding, with a simple “taken” or “snoozed” reply.
5. Measurement & Guardrails
A product can hit its target metric while quietly causing damage somewhere else, which is why this stage insists on both a Primary Outcome Metric and a set of guardrail metrics running alongside it. For Northbridge, the primary metric might be the percentage of reminders marked “taken” within two hours of the scheduled time. Guardrails would include complaint volume, opt-out rate, message delivery reliability, and, given the population, any indication of confusion that could lead to a missed or duplicated dose, a genuine safety concern in a health context. A successful outcome should not hide an unacceptable side effect. Measurement and guardrails are set before the experiment runs, not chosen afterward to justify whatever happened.
6. Experiment Planning
Vague optimism is not a hypothesis. The framework asks teams to state their bet in a fixed format: “We believe [customer] will achieve [outcome] by using [solution/output]. We will consider this supported when [evidence] reaches [threshold] within [time period].” For Northbridge: “We believe patients over 60 managing multiple prescriptions will improve on-time medication adherence by using a daily text reminder. We will consider this supported when the reminder-response adherence rate reaches 70 percent within a four-week pilot.” Pass and fail thresholds are set in advance, in this planning stage, precisely so nobody can move the goalposts once the results are sitting in front of them. This is the point in the sequence where experiment planning turns a belief into something falsifiable.
7. Review, Learn & Iterate
Once an experiment runs, the cycle is Learn, Decide, Change, Retest, and it ends in one of five decisions: Continue, Improve, Pivot, Pause, or Stop. Continue means the evidence supports moving forward as planned. Improve means the direction is right but the execution needs adjustment. Pivot means the underlying assumption needs rethinking. Pause means more information is needed before deciding either way. Stop means the evidence does not support the outcome, and the resources are better spent elsewhere. Say Northbridge’s pilot lands at 55 percent adherence against a 70 percent threshold, but exit interviews reveal that most non-responders simply did not notice the text among other notifications. That is not a Stop signal, it is an Improve signal, pointing toward a different notification channel or a more visually distinct reminder before reviewing, learning, and iterating again.
8. Product Logic & Iteration Loop
Each iteration should be summarized in one sentence that ties the customer, the output, the problem, the outcome, and the evidence together: “For [core customer], this [minimum output] will address [priority problem] and help achieve [desired outcome], measured by [evidence].” For Northbridge’s second iteration: “For patients over 60 managing multiple prescriptions, this redesigned reminder with a distinct sound and a larger on-screen alert will address missed doses caused by notification fatigue and help achieve improved adherence, measured by the percentage of reminders marked taken within two hours.” Keeping this sentence current is a small habit with an outsized effect: it forces the product logic and iteration loop back into a single, checkable claim every time the plan changes, rather than letting the rationale drift silently across meetings.
9. Implementation & Action Plan
Good intentions do not ship anything on their own. Every action in the plan needs to be tied to an outcome it serves, an owner who is accountable for it, a date it is due, any dependency it is blocked on, and the evidence that will show it was actually completed, not just started. For Northbridge, “redesign the reminder alert” is owned by a specific designer, due before the next pilot window opens, dependent on the mobile team’s notification API update, and considered complete only when the new alert is live for all pilot participants and confirmed by a screenshot audit, not merely marked done in a project tracker. This is where an implementation and action plan either holds a team accountable to the outcome or quietly turns back into a generic task list.
10. Product Readiness & Final Decision
Before something moves toward a broader release, it earns one of three verdicts: Ready, meaning the evidence and the guardrails both support moving forward without reservation; Ready with conditions, meaning it can proceed but only alongside specific fixes or monitoring; or Not ready, meaning the evidence does not yet support the outcome claim. For Northbridge, the redesigned reminder might reach “Ready with conditions”: clear to expand from fifteen pilot patients to the full population of eligible patients, on the condition that the opt-out rate is monitored weekly for the first month and that a fallback phone-call option remains available for patients whose adherence does not improve. Product readiness is a formal checkpoint, not an assumption that momentum alone justifies moving forward.
11. Scaling & Sustainability
Scale is earned, not assumed. This stage says a product should expand only after its value has been shown to be repeatable, not merely present in a single favorable pilot. Growth should strengthen the customer outcome, not dilute it. For Northbridge, scaling from fifteen pilot patients to several thousand only makes sense if the adherence gains, the guardrail behavior around opt-outs, and the operational cost of message delivery all hold up at that larger size, and if the team has confirmed the outcome is not an artifact of the unusually high attention a small pilot group naturally receives. Scaling and sustainability is the stage that keeps a genuinely good early result from being generalized past what the evidence actually supports.
12. Product Lifecycle & Exit Strategy
Every product eventually moves through Launch, Growth, Maturity, Decline, and Transition, and the framework treats the last of these as a real design responsibility rather than an afterthought. Responsible retirement includes a plan for customer transition, so people are not abruptly left without something they depend on, clear handling of any data collected along the way, and knowledge transfer so the reasoning behind past decisions is not lost when a team moves on. If Northbridge eventually replaces its text-based reminder with an integrated in-app notification system, the product lifecycle and exit strategy stage requires notifying patients well in advance, migrating their reminder preferences rather than resetting them, and documenting what was learned about notification fatigue so the next team does not have to relearn it from scratch.
Three Considerations That Run Throughout
Three areas do not fit neatly into a single stage because they apply continuously, from the first customer conversation through the final retirement decision. Feasibility and Viability asks whether a team can actually build, run, and afford to sustain what it is proposing, a question worth revisiting at every stage rather than only at the start. Responsible Product Risk matters especially wherever AI is involved: what happens when the model is wrong, who is affected when it is, whether human review is required before an automated decision reaches a customer, and whether the customer can opt out. The NIST AI Risk Management Framework’s four functions, Govern, Map, Measure, Manage, and its emphasis on human-in-the-loop review for consequential decisions, offer a useful structure for thinking this through, particularly relevant for anything resembling Northbridge’s reminder logic if it were ever extended to include automated dosage suggestions. Adoption and Change is the third: a good product can still fail if people do not actually adopt it, regardless of how sound the underlying outcome logic was.
A good product can still fail if adoption fails. Building the right thing and getting people to actually use it are related problems, but they are not the same problem, and solving the first does not automatically solve the second.
Pendo’s 2019 Feature Adoption Report, based on an analysis of roughly 615 software subscriptions, found that close to 80 percent of features were rarely or never used. That is one vendor’s data, not a universal law, but it lines up with a pattern many product leaders recognize: building something is necessary but not sufficient. Teresa Torres’s “Continuous Discovery Habits” argues for weekly, ongoing contact with customers, using tools like Opportunity Solution Trees, precisely so adoption problems surface while there is still time to act on them rather than in a post-mortem.
Prioritization and Discipline
None of the twelve stages tell a team how to choose among competing good ideas once several have survived the earlier filters, which is where established prioritization tools still earn their place inside this framework rather than being replaced by it. RICE, developed at Intercom and associated with Sean McBride, and ICE, associated with Sean Ellis, both give teams a repeatable way to rank options once the outcome case for each has already been made. This framework is upstream of those tools, making sure what feeds a scoring model is a genuine, evidence-backed outcome hypothesis rather than a feature someone liked in a meeting.
It is also worth naming a claim this framework does not lean on. The Standish Group’s CHAOS Report is often cited for statistics about project success and failure rates, but it is a famous, methodologically contested industry claim, and nothing here depends on those numbers being accurate. The argument rests on a more basic observation: measuring output is easier than measuring outcome, and organizations tend to manage what is easy to measure unless someone deliberately builds a structure that measures the harder, more important thing instead.
How to Actually Start
A team does not need to run all twelve stages in full ceremony to get value from this framework. The highest-leverage place to start is the first three stages: name the core customer honestly, even if it means admitting the product currently serves nobody in particular; write down the actual problem and where it sits in the outcome chain, distinguishing what happens immediately from what constitutes real success; and separate what is known from what is merely assumed, because that single habit tends to surface more risk than any other step in the sequence. From there, a Minimum Viable Output and a single Primary Outcome Metric are usually enough to run a first real test.
Downloadable worksheets for each stage, including templates for the assumption log, the hypothesis statement, and the readiness checklist, are available in the resources section of this site. For teams that want direct help applying the framework to a specific product decision, that work is available through consulting.
Key Takeaway
Outcome driven product design is a discipline for deciding what deserves to be built before deciding how to build it well. The twelve stages, from customer segmentation through lifecycle retirement, exist to keep that decision anchored to evidence rather than enthusiasm, especially now that AI has made producing output so cheap that the temptation to skip straight to building is stronger than it has ever been. The framework does not slow a team down for its own sake. It spends effort earlier, where a wrong turn is still cheap to correct, so that speed later is spent on something worth building.
Outcome driven product design reverses the usual order of product work, starting with the customer and the desired outcome and only then deciding what to build, so that speed of execution never substitutes for evidence that the right thing is being built.