Mental model

Metric Auditing

A systematic process for evaluating whether your performance measures actually reflect your goals or incentivize unintended behavior.

Discover

A hospital's safety rating improved dramatically when administrators announced they'd measure success by tracking patient complications. Years later, an investigation revealed something disturbing: surgeons had started avoiding high-risk cases to protect their scores, meaning the sickest patients couldn't find care.

What went wrong?

See how to check whether your metrics measure what matters.

Understand

Understand

Auditing your metrics means checking whether the numbers you track actually reflect what you care about, or if people are just learning to game the system. When you make something measurable and tie rewards to it, people naturally adjust their behavior—which sounds good but can backfire. For example, call center workers told to minimize call length might hang up on customers just to keep their numbers down. Try this: Pick one metric you track and ask what someone would do if they only cared about that number.

Full explanation

Full explanation

How Metric Auditing Works

Metric auditing follows a straightforward process: identify your intended goal, examine what you're actually measuring, then trace how people might respond to those incentives. Start by stating the goal in plain language, not numbers. If your goal is "provide excellent customer support," but you only measure "average call time," you've created a gap between intent and measurement.

The Gaming Mechanism

When metrics determine rewards, people optimize for the metric rather than the goal. This happens through a predictable sequence: First, the metric is introduced with good intentions. Next, people discover which actions most efficiently improve their numbers. Finally, those actions may drift away from the original goal, sometimes dramatically. A sales team rewarded on revenue might push expensive products customers don't need, while teachers evaluated on test scores might focus only on test-taking strategies at the expense of deeper learning.

Warning Signs to Check

Look for three red flags: behaviors that look technically correct but feel wrong, shortcuts that sacrifice quality for speed, or results that improve suspiciously fast without corresponding improvements elsewhere. In healthcare, surgeons avoiding difficult cases to maintain their mortality rates is a classic example—numbers look great, but patient access suffers.

Practical Steps

First, diversify your measurements. No single number captures complex goals. Combine quantitative metrics with qualitative feedback, and measure multiple dimensions of performance rather than one easily-gamed statistic. Second, keep some metrics secret or change them periodically so people can't simply optimize for known targets. Third, use metrics as diagnostic tools rather than sole determinants of rewards—understand what the numbers tell you about the system, but make decisions based on broader judgment.

When Auditing Matters Most

Audit your metrics whenever stakes are high, when you notice suspicious patterns, or when implementing new measurement systems. Pay special attention when metrics replace human judgment entirely, or when people who are measured have influence over the data itself. The more weight a metric carries in decisions about hiring, firing, funding, or bonuses, the more carefully you should audit it for unintended consequences.

Research

Research

Research on metric corruption reveals that when quantitative measures become tied to high-stakes decisions, they inevitably distort behavior. Campbell's Law, formulated by social scientist Donald T. Campbell in 1979, states that "the more any quantitative social indicator is used for social decision-making, the more subject it will be to corruption pressures and the more apt it will be to distort and corrupt the social processes it is intended to monitor"[1]. Economist Charles Goodhart independently observed a similar phenomenon in monetary policy, noting that "when a measure becomes a target, it ceases to be a good measure"[2].

Recent comprehensive research by Manheim (2024) synthesizes these findings and provides practical strategies for building less-corruptible metrics, including diversification, randomization, and post-hoc specification[3]. Jerry Z. Muller's work on "metric fixation" demonstrates how the obsession with measurement has spread across education, healthcare, policing, and business, often replacing professional judgment with bureaucratic compliance[4]. Marilyn Strathern's research on audit cultures in higher education shows how performance ratings can paradoxically undermine the very quality they aim to improve[5].

Empirical studies document these failures across domains. A 2024 UCLA Health study found that a federal patient safety metric (PSI-04) was fundamentally flawed for stroke care, potentially creating incentives to avoid performing lifesaving procedures on the sickest patients[6]. In education, standardized testing pressure has led to "teaching to the test" phenomena where students perform well on exams without developing corresponding understanding or skills.

The research suggests several mitigation strategies: using multiple measures rather than single metrics, keeping some metrics secret to prevent gaming, incorporating qualitative human judgment alongside quantitative measures, and periodically reviewing metrics for unintended consequences. Perhaps most importantly, metrics should serve as diagnostic tools to inform decision-making rather than replace it entirely.

Limitations

Limitations

Metric auditing has important limitations. Not everything worth doing can be easily measured, and some attempts to quantify complex goals may do more harm than good. In creative fields, research-intensive work, or caregiving professions, rigid metrics may miss crucial dimensions of performance. Additionally, some organizations lack the resources or expertise to conduct thorough metric audits, and secret metrics can undermine trust and transparency. There's also a tension between preventing gaming and providing clear expectations—people deserve to know how they'll be evaluated. Finally, metrics may be necessary for accountability in large organizations where personal supervision isn't feasible; abandoning metrics entirely isn't always practical or desirable.

Try it

Synthesize

Choose a pattern from the guide, then pick an action to try with it.

Which pattern stands out?

What will you try?

Choose a pattern above to select an action.

Sources

Sources

Try it

Check your understanding

A software team tracks "lines of code written" as a measure of productivity. Over time, managers notice code becoming more verbose and maintenance costs rising. What's the most likely explanation?

Show the guide's explanation

Answer: Programmers are adding unnecessary code to hit targets

This demonstrates metric gaming in action. When 'lines of code' becomes the target, developers have an incentive to write verbose, redundant code—even though this makes software harder to maintain and contradicts the actual goal of writing efficient, high-quality software. The metric has become decoupled from true productivity.

Which of the following is the most effective strategy for reducing metric manipulation in a high-stakes performance evaluation system?

Show the guide's explanation

Answer: Combine multiple measures with qualitative human judgment

Research shows that diversification—using multiple measures that are harder to game simultaneously—combined with qualitative judgment is more effective than any single strategy alone. Secret metrics can undermine trust, and frequent measurement just accelerates gaming. Human judgment provides a check against metric distortions and captures dimensions of performance that numbers miss.

True or False: If people are successfully 'gaming' your metric, this means the metric calculation is wrong.

Show the guide's explanation

Answer: False

Gaming doesn't mean the calculation is wrong—it means the metric has become decoupled from the goal. The math may be perfect, but people are responding to incentives in ways that undermine the intended outcome. This is the core insight of Campbell's Law and Goodhart's Law: any metric tied to rewards will eventually be optimized in ways that may not serve the original purpose. The solution is usually to reconsider what you're measuring and why, not just fix the formula.

Keep exploring

Find another idea for the decision in front of you.

The complete Reframo library is free to read. Explore another guide whenever you are ready.