Ask a product team why they built a feature, and most will answer in terms of what it does. Ask them what will happen because they built it, and the answers get vaguer fast, often landing directly on something like “it will improve retention” or “it will drive engagement.” Between those two answers there is an enormous, unexamined gap, and that gap is where most product bets quietly fail without anyone noticing until a quarterly review forces the question.
The gap has a name: the outcome chain. It is the sequence of things that have to actually happen, in order, between a customer touching what you built and the long-term result you are hoping for. Skip a link, and you are not managing a product outcome, you are hoping for one.
Why a Chain, Not a Straight Line
Joshua Seiden’s book “Outcomes Over Output” made the case that an outcome is a change in customer behavior, not a change in what a team ships. That distinction matters here because it points at something teams routinely miss: a behavior change does not happen the instant a feature goes live. It happens in stages, and each stage has to hold up before the next one can.
The outcome chain used in this framework has four links:
- Immediate Outcome, the first observable change that happens when a customer uses the output.
- Behavior Change, a new or altered pattern of action the customer repeats.
- User Success, the customer actually achieving something they value.
- Long-Term Impact, the durable, often business-relevant result that follows from many customers succeeding over time.
Each link is a claim, and each claim needs its own evidence. You cannot inspect a long-term impact claim directly on launch day, but you can inspect the immediate outcome and the behavior change almost right away, and that is exactly why the chain is useful. It gives you something to check before the metric you actually care about has had time to move.
To keep this concrete, imagine a hypothetical budgeting app. Its team has just shipped a new feature: automatic categorization of bank transactions, so a user no longer has to manually tag every purchase as “groceries” or “dining out.” The team’s internal pitch deck says this feature “will increase retention.” Let’s walk that claim through all four links and see what is actually being assumed.
Link One: Immediate Outcome
The immediate outcome is the smallest, most direct thing that changes because the output exists. It is not a business metric. It is closer to a moment: does the customer notice something different, and does it work the way it was meant to?
For the budgeting app, the immediate outcome is something like: the user opens the app, sees their transactions already sorted into categories, and does not have to tag them manually. That is observable within a single session. You can watch it happen, or measure it through event logs, the day the feature ships.
This is the link teams get right most often, because it is the closest one to the output itself. It is also the least interesting one on its own. A user seeing pre-sorted transactions is nice, but it is not yet a reason for the business to care.
Link Two: Behavior Change
This is the link most roadmaps quietly skip, and it is the one this article spends the most time on, because it is both the most commonly missing and the fastest to test.
Behavior change means the customer starts doing something differently, repeatedly, not just once. For the budgeting app, a plausible behavior change hypothesis is: users who get automatic categorization start checking their spending breakdown weekly, instead of opening the app only when they are worried about their balance. Another candidate hypothesis: users start adjusting a discretionary spending category, like dining out, after seeing it visualized clearly for a few consecutive weeks.
Notice that neither of these is the same claim as “users see categorized transactions.” Seeing a feature is not a behavior change. A behavior change is a new pattern that persists past the first encounter with the output. It is the difference between a customer trying a feature once out of curiosity and a customer building a new habit around it.
A shipped feature is a fact about your product. A behavior change is a fact about your customer’s life. Only one of those two facts is worth reporting to leadership as progress.
This link is where the needs, pains, and desired outcomes that motivated the feature in the first place either get validated or exposed as wishful thinking. If the team’s real desired outcome was “users make better spending decisions,” then a behavior change involving actually looking at spending data on a recurring basis is a prerequisite. If that recurring look-in never happens, nothing downstream of it will happen either, no matter how good the categorization algorithm is.
Link Three: User Success
User success is the customer actually getting the value they were after, not just performing a new behavior, but benefiting from it. For the budgeting app, this might be: the user’s discretionary spending in a tracked category goes down over a month, or the user builds a small savings buffer they did not have before.
This is a meaningfully different claim from behavior change. A user can check their spending breakdown every week (behavior change) and still overspend every month (no user success), if the visualization alone does not change their decisions at the point of purchase. User success requires the behavior to actually connect to the underlying need, which is why the assumptions and evidence behind this link deserve to be written down and challenged, not assumed to follow automatically from the behavior above it.
This is also where problem strength matters. If overspending was a low-frequency, low-severity annoyance for most users rather than an urgent, costly pain, a categorization feature might produce a behavior change without ever producing a meaningful success, because there was not enough pressure behind the original problem to begin with.
Link Four: Long-Term Impact
Long-term impact is the durable result that shows up after enough customers experience user success over enough time. For the budgeting app, this is where “increased retention” actually belongs, along with things like higher subscription renewal, more referrals, or a reputation as a tool that genuinely helps people spend less anxiously.
This link is real, and it is usually the one that matters most to the business. It is also the slowest to observe, the easiest to confound with unrelated factors like seasonality or a pricing change, and the one that tells you the least about what to fix if it does not materialize. If retention does not improve six months after the categorization feature ships, that fact alone gives the team almost nothing actionable. Which link broke? Was there no user success? No behavior change? Did the immediate outcome even land the way it was designed to? Without the chain, you are debugging a system with the instrumentation ripped out.
The Mistake: Jumping Straight to Impact
Here is the pattern that shows up again and again in product reviews: a team ships an output, and the very next slide claims a long-term impact, with nothing named in between. “We shipped auto-categorization. This will increase retention.” No mention of what new behavior the feature is supposed to produce, and no mention of what success looks like for the customer along the way.
This habit is understandable. Long-term impact metrics, retention, revenue, lifetime value, are the ones executives ask about, so they are the ones that get put on slides. But naming only the first link (the output) and the last link (the impact) skips the two links that are actually within the team’s power to observe, influence, and correct in the near term. It also means that if the long-term number does not move, there is no diagnostic trail to follow. The team is left re-litigating the entire feature from scratch instead of checking a specific, falsifiable link in the chain.
Melissa Perri’s description of the “build trap” is relevant here: teams measure themselves by what they ship rather than by what changes for the customer. Naming only the output and the hoped-for impact, with the middle left blank, is the build trap wearing an outcome-shaped costume. It sounds like outcome thinking because it mentions retention, but there is no actual mechanism connecting the feature to that number.
Finding the Weakest Link in Your Chain
Most teams, when they actually write out their outcome chain link by link, find that one link is far more speculative than the others. Usually it is not the immediate outcome, since that is easy to verify by watching a user session. It is also, somewhat counterintuitively, often not the long-term impact link either, because teams frequently have historical data or industry benchmarks that make a rough impact estimate plausible.
The weakest, most speculative link is almost always behavior change. Teams can describe what the output does. They can describe the business result they want. What they frequently cannot describe with any confidence is the specific, repeatable action the customer will take differently because the output exists. That gap is worth taking seriously, because it is also the cheapest one to close.
Behavior change is the fastest link in the entire chain to observe. You do not need months of retention data or a full success-metric pipeline. You need usage logs, session recordings, or a short qualitative check-in a week or two after a small group of customers gets access to a minimum viable output. Did the new pattern of behavior actually show up, or didn’t it?
This maps directly onto the framework’s experiment structure: “We believe [core customer] will [check their spending breakdown weekly] by using [automatic categorization]. We will consider this supported when [at least 40 percent of the pilot group open the breakdown view at least once a week] within [three weeks].” That is a testable claim about the second link, and it can be answered in weeks, long before anyone has enough data to say anything credible about retention.
If the behavior change does not show up, there is no reason to wait around for a long-term impact number to confirm what the second link already told you. The right move at that point is not to keep shipping adjacent features and hoping one of them sticks. It is to go back to the assumptions behind the behavior change hypothesis and ask what was actually wrong: wrong customer segment, wrong problem, wrong solution, or a real problem paired with a solution that did not remove enough friction to change the habit.
Teams that want a structured way to walk an existing roadmap through this kind of link-by-link audit, rather than discovering the gap after a launch has already underdelivered, may find it useful to work through it with outside eyes; that is one of the things consulting support on this framework is built to help with.
Key Takeaway
An outcome chain connects an immediate outcome to a behavior change, a behavior change to user success, and user success to long-term impact, and none of those links can be safely assumed just because the one before it happened. The single most common failure is naming only the output shipped and the long-term impact hoped for, while leaving the behavior change in between unnamed and untested. Since behavior change is usually the weakest and fastest-to-observe link in the chain, test it first, before spending months waiting on a retention or revenue number to tell you what a two-week usage check could have told you already.
An outcome chain has four links, immediate outcome, behavior change, user success, and long-term impact, and skipping the behavior change link to jump straight from a shipped feature to a retention claim is the most common way teams fool themselves. Test the behavior change link first, because it is the fastest one to observe.