The Measurement Trap

The moment we begin using a metric to measure performance, it starts to drift from its original purpose. University entrance exams in Korea and many other countries are an easy example. They were presumably designed to select students with the academic ability they would need. Yet we all know that academic ability and an exam score are not the same thing. Once the exam itself becomes the goal, being good at taking the exam matters more than the ability it was meant to measure.

Lines of code and story points create the same problem in software organizations. If lines of code become a target, short and clear code is penalized. If story points become a measure of performance, estimates turn into a scorecard. They were supposed to be a language for planning. Evaluation is ultimately qualitative work: someone has to invest the effort to read the context. Removing judgment may look fair, but removing the context along with it only distorts the evaluation.

In the stock market, funds are evaluated by their alpha, the amount by which their returns exceed a market index. Yet that number conceals a great deal. We also need to ask how much risk was taken, how much of the result came from luck, and whether long-term gains were sacrificed to produce short-term alpha. A number can summarize the outcome, but it cannot fully explain the process.

Evaluating people is no different. Managers are not always deeply committed to doing evaluations well. For the sake of the example, though, suppose a manager sincerely wants to evaluate everyone as fairly as possible. The dilemma remains. No measurement system can fully capture a person’s qualitative contribution.

Imagine a designer whose particular strength is artwork. Their work raises a landing page’s conversion rate by 3%, the best result the company has ever seen. The evaluation form, however, divides the abilities expected of a designer into separate scores. On a form like this, someone with one exceptional strength is at a disadvantage. A person who is unremarkable but competent in every category is more likely to receive the higher score.

The qualitative categories make things even more complicated. This designer pushes hard for the work they believe will produce the best result and sometimes accepts the conflict that follows. It does not work every time, but it usually leads to a good outcome. They are positive and enthusiastic, but they have a child and cannot stay late at the office. They continue working at home, where that effort is invisible. Their score for visible enthusiasm ends up lower than that of a new graduate.

The designer receives a B-; the new hire receives an A. Luck may have played a part in the conversion-rate increase, and no single result can explain every contribution. The underlying problem still remains: no metric can hold the whole of one person’s contribution. A manager with strong biases will make a qualitative evaluation even less fair. But does trusting the evaluation form make those biases disappear? It only hides them behind scores and coefficients.

If we want to evaluate everyone fairly, we should spend less time adding scores and adjusting coefficients and more time discussing specific contributions in their actual context. Evaluation is not a responsibility that numbers can take over. People have to read the context and evaluate what other people contributed.