← All articles Product Metrics

How to Know Whether a Product Is Actually Working

Releases shipped and logins counted are not proof a product works. Here is how to measure product success with a Primary Outcome Metric and guardrails.

outcome-metricsguardrailsproduct-successmeasurement

A product team reports a good quarter. Three releases shipped on schedule. Login counts are up eleven percent. Page views on the new dashboard are climbing steadily. Everyone in the room nods, the slide gets a green checkmark, and the meeting moves on. Then someone, usually a person newer to the team or a customer-facing person invited by accident, asks a plain question: did any of this make a customer’s actual problem better? Silence. Nobody measured that. They measured what the team did. Nobody measured what happened to the customer because of it.

This gap is not rare, and it is not a sign of a careless team. It is what happens by default when a product’s dashboard is built around what is easy to count rather than around what the product was supposed to change. Learning how to measure product success starts with a distinction that sounds obvious once stated and gets ignored constantly in practice: output metrics describe what a team made, outcome metrics describe what got better for the customer because they made it.

Output Metrics Feel Like Progress Because They Are Easy to Count

Releases shipped, story points closed, logins per week, page views, session counts, features launched. All of these are real numbers, and none of them are fake in the sense of being made up. The problem is what they actually measure. They measure activity and exposure, not benefit. A login tells you someone opened the product. It does not tell you they got what they came for. A page view tells you a screen rendered. It does not tell you the person looking at it understood anything, decided anything, or left better off than when they arrived.

Output metrics became the default reporting language for a boring reason: they are cheap to instrument, they update in real time, and they travel well through a slide deck. “We shipped four features and logins are up” is a sentence a VP can repeat verbatim in the next meeting up the chain. “Customers are resolving their billing disputes forty percent faster now” requires someone to have defined, measured, and tracked that specific change deliberately. It does not happen automatically the way a login counter does.

Output is what we make. Outcome is what becomes better because we made it.

That distinction is the core reversal behind outcome-first product thinking, and it is worth sitting with because most existing dashboards were not built with it in mind. A team can hit every output number on its scorecard and still be unable to answer whether the product is achieving anything customers care about.

Why This Gets Worse, Not Better, With More Output

It is tempting to think that more releases and more usage naturally add up to more value over time. Pendo’s 2019 Feature Adoption Report, based on an analysis of roughly 615 software subscriptions, found that around 80 percent of features in the products it studied were rarely or never used after launch. That is one vendor’s dataset, not a universal law, but it points at something real: shipping volume and customer benefit are not the same curve. A team can ship faster and faster while the share of that output that actually changes anything for a customer stays flat or even shrinks, because velocity was never the thing standing between the team and customer value in the first place.

What an Outcome Metric Actually Looks Like

An outcome metric describes a change in the customer’s world, not a change in the product’s activity log. Time saved completing a task. A problem resolved without escalation. Confidence increased before a decision. An error rate reduced. A behavior changed, such as a patient completing a prescribed exercise routine instead of abandoning it after week one. These are the kinds of changes the Outcome Chain is built to trace: an immediate outcome leads to a behavior change, which leads to user success, which eventually compounds into long-term impact. Skipping straight to “long-term impact” claims in a launch review, before the immediate outcome and behavior change have even been measured, is how teams end up defending a metric they cannot actually explain.

Every credible outcome metric needs two anchors before it means anything: a baseline and a target.

Baseline is the honest, current-state number, measured before the change ships, using the same definition and the same data source the team intends to use afterward. If a team cannot state its baseline, it has no way of knowing whether anything moved, and every post-launch number gets argued about in a vacuum.

Target is a specific, time-bound threshold the team believes represents a meaningful improvement, not an arbitrary round number. A target of “reduce it by 25 percent within eight weeks” can be argued with, tested, and eventually judged. A target of “make it better” cannot. This is the same discipline behind a well-formed experiment: “we believe [customer] will achieve [outcome] by using [output]. We will consider this supported when [evidence] reaches [threshold] within [time period].” A Primary Outcome Metric is that same threshold logic, applied at the level of the whole product or feature rather than a single test.

Setting both numbers before launch, not after, matters for a reason that has nothing to do with rigor for its own sake. A baseline set after the fact is really just a guess dressed up as data, and a target set after seeing the results is not a target, it is a rationalization.

Guardrails: What the Outcome Metric Alone Will Not Tell You

Here is the part that gets skipped most often, and it is the part that causes the most damage when it is skipped. A product can hit its outcome target cleanly and still be doing real harm somewhere else that nobody was watching.

Picture a scheduling feature designed to reduce the time customers spend booking an appointment. The team ships it, the average booking time drops from six minutes to two, the outcome target is hit ahead of schedule, and the launch is called a success. Three months later, support tickets related to double-bookings and missed appointment reminders have tripled. Nobody connected the two, because nobody was tracking support ticket volume as part of that launch. The feature made booking faster by quietly skipping a confirmation step that used to catch scheduling conflicts. The Primary Outcome Metric looked great. The customer experience, taken as a whole, got worse.

A successful outcome should not hide an unacceptable side effect.

This is why guardrail metrics have to be defined before launch, alongside the Primary Outcome Metric, not added afterward once something has already gone wrong. Typical guardrails include complaint volume, safety incidents, error rate, cost per transaction, privacy exposure, exclusion of a customer segment who cannot use the new flow, support burden, reliability, and in health-adjacent products, potential for human harm. None of these are meant to be optimized. They are meant to be watched, with a pre-agreed threshold that triggers a pause or a rollback if crossed. The measurement and guardrails stage exists because a single metric, however well chosen, is a narrow window onto a much wider system.

One Primary Outcome Metric, Not a Dozen

There is a temptation to guard against every conceivable risk by tracking fifteen metrics at once. In practice this produces the same blindness as tracking none, because a dashboard with fifteen equally weighted numbers gives a team no way to know which one actually matters this month. The discipline is one Primary Outcome Metric, the single number that best represents whether the priority customer outcome is being achieved, plus a small, deliberately short list of guardrails that exist to catch the side effects the primary metric cannot see. Everything else is background context, not part of the scorecard that decides whether the product is working.

A Worked Example: A Medication Reminder Feature

To make this concrete, consider a hypothetical feature: a smart medication reminder inside a digital health app, intended to help patients with a chronic condition take a daily prescribed medication on schedule rather than missing doses. The core customer is patients who have been prescribed a daily medication and have a documented history of inconsistent adherence. The desired outcome is not “the app sends more notifications.” It is “patients take their medication on the days they are supposed to.”

MetricTypeBaselineTargetData SourceOwnerReview Period
Weekly medication adherence ratePrimary Outcome61% of prescribed doses taken on time78% within 12 weeks of launchApp check-in logs cross-referenced with pharmacy refill dataProduct manager, adherence featureBiweekly
Reminder opt-out rateGuardrailNot applicable (new feature)Stay below 15%In-app settings analyticsProduct manager, adherence featureBiweekly
Support tickets tagged “reminder issue”Guardrail4 per week (baseline period, prior notification system)Stay below 8 per weekSupport ticketing systemSupport operations leadWeekly
Notification fatigue complaintsGuardrailNot applicable (new feature)Stay below 2% of active users per monthIn-app feedback form, tagged complaintsProduct manager, adherence featureMonthly
Missed-dose escalation accuracyGuardrailNot applicable (new feature)95% of true missed-dose alerts confirmed accurate by care team reviewCare team review sample, 50 alerts per monthClinical operations leadMonthly

Notice what this table is doing. The Primary Outcome Metric, adherence rate, is the one number that answers “is this feature achieving the outcome it exists for.” Every guardrail exists to answer a different question: is the feature achieving that outcome by annoying people into disabling it, by overwhelming support with a new category of ticket, by producing false alarms that erode trust in the care team, or by generating noise nobody wanted. If adherence climbs to 80 percent while opt-outs spike to 30 percent, the honest read is not “we succeeded.” It is “we found a way to get a good number that customers are actively escaping,” and the guardrail is what catches that before it becomes a pattern serious enough to require a rebuild.

Ownership and review cadence are not administrative details here. A metric with no named owner tends not to get reviewed, and a metric with no set review period tends to get checked only when something has already broken.

The Discipline This Requires

None of this is complicated in concept. It is demanding in practice because it requires defining success and failure conditions before anyone knows how the launch will actually go, when it would be easier to wait, see what numbers move, and build a narrative around whichever ones look good. Committing to a baseline, a target, and a guardrail list before launch removes that flexibility on purpose. It is the same discipline behind the Product Logic Statement: for a defined core customer, a defined minimum output addresses a defined priority problem and helps achieve a defined outcome, measured by defined evidence. “Measured by” is not a formality tacked onto the end of that sentence. It is the part that eventually tells the team whether the rest of it was true.

Output metrics are already sitting in the analytics dashboard the day a feature ships. Outcome metrics have to be deliberately designed, instrumented, and baselined ahead of time. That extra work is what separates a team that can say, with evidence, “this made things better for customers,” from a team that can only say “we shipped it, and usage went up.”

Key Takeaway

Knowing how to measure product success requires separating what a team produced, releases, logins, page views, from what changed for the customer because of it. A Primary Outcome Metric needs an honest baseline and a specific, time-bound target set before launch, and it needs a small set of guardrail metrics, watched on the same cadence, so a good headline number cannot quietly hide complaints, errors, cost overruns, or harm building up somewhere else in the system.

Key takeaway

Knowing how to measure product success starts with separating what the team produced from what changed for the customer, and it stays honest only when a small set of guardrail metrics is defined before launch, not after something goes wrong.

Bring this to a real decision

The Outcome-Driven Product Design Framework, created by Dr. Mashiur Rahman, hosted under ComingTechs Advisory.

Book a strategy session → Read the framework
Related articles
← All articles