A team ships eighteen features in a quarter. The sprint reports look great. Velocity is up. The roadmap slide is full of checkmarks. Then someone asks the harder question: how many of those eighteen features changed what a customer actually does, or how a customer actually feels about the product? Often nobody in the room can answer, because nobody measured that. They measured shipping. They did not measure the thing shipping was supposed to cause.
This is not a story about a lazy team or a careless manager. It is a story about a measurement habit that made sense for a while and then quietly stopped making sense, while everyone kept using it anyway.
Why Features Shipped Became the Default Scorecard
Counting features shipped, tickets closed, or story points completed did not become the default team scorecard by accident. It became default because it solved three real organizational problems.
First, it is visible. A shipped feature is a discrete, demonstrable event. You can point to a release note, a changelog entry, a demo. Compare that to “customer trust increased,” which is fuzzy, slower to observe, and harder to put on a slide.
Second, it is easy to report upward. A VP asking “what did engineering do this quarter” wants a number, and “we closed 340 story points” is a number. It travels well through layers of management precisely because it requires no interpretation.
Third, it is easy to compare sprint over sprint. Velocity as a planning input, estimating how much a team can commit to next sprint, is a legitimate and narrow use of story points. The trouble starts when that planning number quietly gets promoted into a performance and productivity metric, used to judge whether the team is doing good work rather than just how much work capacity it has.
None of that makes the underlying premise true. It only makes the premise convenient. And convenience is a poor substitute for correctness when the number in question is being used to decide whether a team, or a roadmap, is actually working.
A team can be extremely busy and extremely unproductive at the same time, if busy means “shipped output” and productive means “created customer outcomes.”
Marty Cagan and the SVPG community have described the resulting pattern as the “feature factory”: a team structure optimized for throughput of requested output, with no real mechanism for asking whether that output was the right thing to build. Melissa Perri named this same failure mode the “build trap” in Escaping the Build Trap, the reflexive assumption that more shipped equals more value delivered. Both descriptions point at the same root cause: the scorecard measures the make step, not the change step. Output is what we make. Outcome is what becomes better because we made it.
The Evidence That Shipping Is Not the Same as Value
This is not just a theoretical objection. Pendo’s 2019 Feature Adoption Report, an analysis of roughly 615 subscription-based software products, found that around 80 percent of features in the products it studied were rarely or never used by customers after release. That number comes from one vendor’s dataset, not a universal law of software, and it should be read as directional evidence rather than settled fact. But directionally it lines up with what most experienced product and engineering leaders have seen firsthand: a meaningful share of what gets built ships, sits in the product, and never gets touched.
If a large share of shipped features are barely used, then a scorecard built entirely on “features shipped” is, by construction, rewarding a lot of activity that produces no measurable customer benefit. The team hits its velocity target and the customer experience does not move. Worse, every unused feature still carries an ongoing cost: it has to be maintained, documented, tested against regressions, and explained to new customers, whether or not anyone benefits from it.
CB Insights’ review of startup post-mortems found that “no market need” is the most commonly cited reason startups fail, cited in roughly 42 percent of cases they analyzed. Teams were not failing to ship. They were shipping things nobody needed, on schedule, sprint after sprint, right up until the money ran out.
The AI era makes this risk sharper, not smaller. AI coding assistants and generative tools have made it materially cheaper and faster to produce software, prototypes, and content. GitClear’s 2025 research on code quality, tracking churn and duplication trends from 2020 to 2024 as AI-assisted coding scaled, found rising code churn and more duplicated code alongside that increase in output. Producing more code faster did not automatically produce better code, and it certainly does not automatically produce better decisions about what to build. Velocity without direction creates waste faster. The strategic question was never really “can we build it.” It has always been “should we build it, for whom, and why.” AI just made it cheaper to avoid asking that question and still look busy.
What an Outcome-Based Scorecard Looks Like in Practice
The alternative is not “stop measuring” or “measure nothing objective.” It is measuring a different thing: whether the output changed a customer’s situation, not whether the output shipped.
In the Outcome-Driven Product Design Framework, the measurement stage is built around one Primary Outcome Metric plus a small set of guardrail metrics, rather than a long dashboard of everything that can be counted.
The Primary Outcome Metric answers a single question: is the priority customer outcome actually improving? Concretely, and hypothetically, imagine a team building a claims-status feature for a health benefits app. A features-shipped scorecard says: “we launched the claims-status screen this sprint.” An outcome scorecard says: “the percentage of members who can find their claim status without contacting support went from 41 percent to 68 percent within six weeks of launch.” One is an event. The other is a measured change in customer behavior, closer to what Joshua Seiden describes in Outcomes Over Output as a change in customer behavior that drives business results.
Guardrail metrics sit alongside the Primary Outcome Metric to catch the failure mode where a number looks good but something else quietly got worse. For the same claims-status example, guardrails might include support ticket volume for claims questions, error rate in the status data shown, and complaint rate about the new screen. A successful outcome should not hide an unacceptable side effect. If self-service usage climbs but support complaints about wrong information also climb, that is not a win, it is a new problem wearing the costume of a win.
A practical outcome scorecard, reviewed on the same cadence a features-shipped report used to be reviewed, might look like this:
- Primary Outcome Metric: the one number that says whether the customer’s situation improved (for example, self-service resolution rate, time-to-first-value, task completion rate, error rate reduction).
- Guardrail metrics: the two to four numbers that would reveal an unacceptable side effect (support burden, complaint rate, safety or error incidents, exclusion of a customer segment, reliability).
- Adoption signal: whether customers are actually using the output at all, not just whether it exists. Adoption, including onboarding friction and behavior change, is part of the outcome, not a marketing afterthought tacked on later.
- Decision: given the above, the team’s explicit call, continue, improve, pivot, pause, or stop, on the current output.
That last row matters more than it looks. A features-shipped scorecard has no decision step built in; it just accumulates. An outcome scorecard forces a decision every review cycle, which is a very different discipline for a team to operate under.
Addressing the Real Management Worry
The honest objection from a manager or executive is fair and deserves a direct answer, not a dismissal: “If we stop counting features shipped, how do we know the team is being productive? How do I know they’re not just sitting there?”
The answer is not to stop measuring team activity internally. Sprint velocity, cycle time, and story points completed remain useful as internal planning and forecasting tools, the same way an engineering team tracks build times or deployment frequency. The change is what gets reported upward as evidence of value, and what gets used to judge whether the roadmap itself is working.
A team can report both, cleanly separated: “We completed 34 story points this sprint” is a capacity and planning fact, useful for the team and its immediate lead. “The self-service resolution rate moved from 41 percent to 68 percent, support tickets for this issue dropped 22 percent, and no guardrail moved in the wrong direction” is an outcome fact, useful for anyone deciding whether this initiative deserves more investment. The second statement is not vaguer than the first. It is a real number tied to a real customer behavior, and in most cases it is easier to defend in front of a skeptical executive than a story-point count, because it answers the question executives actually care about: is this working.
This also does not require abandoning agile ceremonies or estimation. It requires adding a layer on top of them: at the point where a feature is planned, it should already carry a hypothesis of the form used in the framework’s experiment planning stage, “we believe this customer will achieve this outcome by using this output, and we will consider it supported when this evidence reaches this threshold within this time period.” That hypothesis is what turns “we shipped it” into “we can now report whether it worked,” on a schedule the business already trusts, sprint reviews, quarterly business reviews, board updates.
Teresa Torres’ continuous discovery practice, weekly structured contact with customers, is one concrete way teams generate the qualitative half of this picture alongside the quantitative Primary Outcome Metric, so the number is not the only signal in the room.
Key Takeaway
Features shipped and story points completed are legitimate internal planning signals, but they were never evidence that customers are better off, and Pendo’s finding that roughly 80 percent of features in its surveyed products go rarely or never used shows how far that gap can stretch. A credible alternative already exists and does not require abandoning agile delivery: keep the internal velocity metrics for planning, and report upward on one Primary Outcome Metric plus a small set of guardrails, reviewed on the same cadence leadership already trusts. That single change turns “we were busy” into “here is what got better, and here is what we watched to make sure nothing got worse.”
Features shipped and story points completed measure how busy a team was, not whether customers are better off. A credible alternative exists: one Primary Outcome Metric plus guardrails, reviewed on the same cadence leadership already trusts.