A metric can be useful for observing a system and harmful when it becomes the only thing the system is rewarded for maximising. Once people optimise the number, they may discover ways to improve it without improving the underlying goal—or while damaging something the metric does not measure.
This family of problems is often associated with Goodhart’s law: when a measure becomes a target, it can cease to be a good measure. The practical lesson is to keep the real objective visible and treat metrics as evidence about it, not as replacements for it.
The student who maximises completed questions
Imagine Ryan’s revision dashboard rewards the number of questions completed each day. At first, completion count is useful because he has been avoiding practice.
Ryan learns how to raise the score: choose easier questions, skip written explanations and avoid difficult mixed problems that take longer. His daily count rises sharply while examination transfer stagnates.
The metric did not become mathematically wrong. Its relationship with the real goal changed after behaviour adapted to the target.
A proxy is not the objective
Question count is a proxy for productive practice. Marks are a proxy for demonstrated performance under a particular assessment. Attendance is a proxy for presence, not learning. Website clicks are a proxy for attention, not necessarily usefulness.
Proxies are unavoidable and valuable. The mistake is forgetting the gap between what is measured and what is actually valued.
Write both: “Goal: transferable examination capability. Metric: independently completed mixed questions with verified correction.” The wording makes it harder for raw volume to become the mission.
Optimisation reveals loopholes
A metric may correlate well with the goal when nobody is deliberately optimising it. Once rewards depend on the metric, users search for the cheapest way to improve the score.
If a tutoring programme rewards tutors only for worksheet completion, time may shift from explanation to answer production. If a project rewards source count, students may add weak references rather than improve evidence quality.
The existing source-count guide is a direct example: increasing the visible count does not guarantee independent or relevant evidence.
One metric can push costs into an unmeasured dimension
Optimising speed can reduce accuracy. Optimising short-term marks can reduce exploration. Optimising utilisation can remove spare capacity needed for disruptions. Optimising word count can create repetitive writing.
These are not universal trade-offs; good systems can improve several dimensions together. The warning is to inspect what the chosen metric makes invisible.
Before setting a target, ask what behaviour could improve this number while making the actual objective worse. That adversarial question often reveals the missing guardrail.
Adding more metrics can create a new game
A common repair is to create a dashboard with ten targets. This can help if the measures represent genuinely important dimensions. It can also create a complicated scoring game where users optimise the formula rather than the work.
Keep a small set of complementary measures and retain direct review of the underlying outcome. Metrics should triangulate reality, not bury it under arithmetic.
Where possible, include measures that are difficult to improve without real progress: fresh transfer tasks, independent verification or user outcomes rather than only activity counts.
Targets can change the population being measured
Suppose a programme is rewarded for its pass rate. One route to improvement is better teaching. Another is accepting only students already likely to pass.
The reported percentage rises, but the population has changed. This connects to aggregation and composition effects: a metric can move because who enters the denominator changed.
Track access and entry criteria when the objective includes serving a broader population. Otherwise the system may improve the measured outcome by avoiding difficult cases.
Do not punish people for revealing problems
If a team is judged only by the number of reported errors, it may reduce reports instead of errors. A healthy system distinguishes discovering a defect from causing it.
In learning, a detailed error log can initially make performance look worse because more weaknesses become visible. That can be progress in diagnosis.
Rewarding absence of reported problems can therefore make the system blind. Pair outcome measures with mechanisms that encourage honest detection and correction.
Use metrics diagnostically, not ceremonially
A metric should trigger a question. If completion volume rises while transfer does not, ask what kind of questions are being selected. If average marks rise while harder tasks disappear, inspect the task mix.
Do not celebrate movement simply because the dashboard is green. Equally, do not dismiss a metric whenever it produces an inconvenient result. Investigate its relationship to the goal.
The value of measurement lies in better decisions, not in producing a number that can be displayed every week.
The repair routine
Write the actual objective without using the metric’s name. Then write what the metric measures directly. Identify at least one way the metric could improve while the objective stays unchanged or worsens.
Add a complementary check for that failure mode. Review changes in population, definitions and task difficulty. Keep direct samples of the underlying work so the metric can be audited against reality.
Finally, revisit the metric after people have adapted to it. A proxy that worked before incentives changed may need redesign once optimisation behaviour becomes visible.
For data interpretation, continue through the Mathematics Article Directory. For study systems, use Why More Practice Can Make the Wrong Method Harder to Fix.
Measure what helps you see the goal. The moment the number becomes more important than the thing it was meant to represent, the system has started working for the dashboard instead of the outcome.
