Define What Good Means Before You Pick Metrics
Choosing Metrics That Actually Align
Companion to The Code Takes Care of Itself · Updated 2026-09-30
This is the full playbook from “Excellence Starts with Standards Specific Enough to Check.” The book prints a shorter version at the end of that chapter, and this page keeps every step and every example.
Use this exercise, Choosing Metrics That Actually Align, if you tell people your team is “high performing” but can’t say what that means in specific, checkable terms. Do the steps in order. Go straight to step three and you get a metrics dashboard nobody trusts.
- Identify your cross-functional stakeholders and the scoreboard each one reads. List every function whose success depends on engineering output: executive leadership, sales, product and growth, customer success, and compliance and legal if your industry needs it. For each function, write down what it counts as a win, in their words, not yours. Don’t write “faster releases.” Write their version: “features shipped in time for the deals we’re trying to close this quarter.”
- Choose one to three metrics for each stakeholder group, and make each one translate engineering activity into that group’s language. Don’t invent a new list; take the candidates from the six-benchmark map in “Excellence Starts with Standards Specific Enough to Check,” which pairs each metric with the stakeholder who reads it (Table 4.1 in that chapter). Adjust to your context, and don’t track everything at once. Nobody checks a dashboard with thirty metrics on it.
- Separate the coaching metrics from the scoreboard metrics. Write the separation down, and tell the team about it. Coaching metrics (individual code quality, pull request times, skill development) belong only in one-on-one growth conversations. Keep them out of company dashboards and stack-ranked reviews. Scoreboard metrics (uptime, NPS, retention, time-to-market) describe outcomes rather than individuals, so you can show them widely. If someone discovers later that you used coaching data in a performance review, the program loses the team’s trust in an afternoon.
- Set a review cadence for each metric, matched to how fast that metric moves rather than to a calendar default. For instance, review incident response and MTTR after each significant event, not only each quarter. Review retention and NPS monthly or quarterly, because they move slowly and a weekly review just produces noise. Review skill development in the one-on-one cadence you already have. Don’t create a separate meeting.
- Every two quarters, review the full set of metrics and ask what each one is causing people to optimize. People will game almost any metric you track long enough, usually without intending to. Attention simply shifts toward whatever the organization measures and rewards.
Economists call this Goodhart’s law, after Charles Goodhart’s 1975 observation about UK monetary policy: “any observed statistical regularity will tend to collapse once pressure is placed upon it for control purposes.” The version people usually quote is “When a measure becomes a target, it ceases to be a good measure.” That’s anthropologist Marilyn Strathern’s 1997 rephrasing.
Both point to the same failure. Suppose you track how long pull requests wait for review, and that number starts falling. That may be good news: reviewers are responding faster and work is moving through the system with less friction. Or it may mean reviewers have learned that the fastest way to improve the metric is to approve changes without reading them carefully. The number improved, but the behavior you actually cared about got worse. Once a metric starts producing that kind of distortion, it has stopped measuring the standard you intended. Retire it, redesign it, or pair it with another measure that makes the tradeoff visible.
Try this with AI.
“Here’s the metric set my team runs on. Play an ambitious engineer who wants to look excellent without doing better work, and tell me exactly how you’d move each number. Then rank those games by how long it would take me to notice.”
The exercise exists to force an explicit answer to a question most CTOs have never written down: when someone outside engineering asks what excellent looks like on this team, can you answer in one sentence per stakeholder, and would your team recognize the answer as true? One more question belongs beside it, from “Build the Vision on Your Company’s Native Alpha”: does anything you measure tell you whether the company’s advantage is growing?
All of the playbooks are listed on the playbooks page.
