Why Optimising One Metric Can Make the Whole System Worse | Goodhart’s Law and Metric Gaming

A metric can be useful for observing a system and harmful when it becomes the only thing the system is rewarded for maximising. Once people optimise the number, they may discover ways to improve it without improving the underlying goal—or while damaging something the metric does not measure.

This family of problems is often associated with Goodhart’s law: when a measure becomes a target, it can cease to be a good measure. The practical lesson is to keep the real objective visible and treat metrics as evidence about it, not as replacements for it.

The student who maximises completed questions

Imagine Ryan’s revision dashboard rewards the number of questions completed each day. At first, completion count is useful because he has been avoiding practice.

Ryan learns how to raise the score: choose easier questions, skip written explanations and avoid difficult mixed problems that take longer. His daily count rises sharply while examination transfer stagnates.

The metric did not become mathematically wrong. Its relationship with the real goal changed after behaviour adapted to the target.

A proxy is not the objective

Question count is a proxy for productive practice. Marks are a proxy for demonstrated performance under a particular assessment. Attendance is a proxy for presence, not learning. Website clicks are a proxy for attention, not necessarily usefulness.

Proxies are unavoidable and valuable. The mistake is forgetting the gap between what is measured and what is actually valued.

Write both: “Goal: transferable examination capability. Metric: independently completed mixed questions with verified correction.” The wording makes it harder for raw volume to become the mission.

Optimisation reveals loopholes

A metric may correlate well with the goal when nobody is deliberately optimising it. Once rewards depend on the metric, users search for the cheapest way to improve the score.

If a tutoring programme rewards tutors only for worksheet completion, time may shift from explanation to answer production. If a project rewards source count, students may add weak references rather than improve evidence quality.

The existing source-count guide is a direct example: increasing the visible count does not guarantee independent or relevant evidence.

One metric can push costs into an unmeasured dimension

Optimising speed can reduce accuracy. Optimising short-term marks can reduce exploration. Optimising utilisation can remove spare capacity needed for disruptions. Optimising word count can create repetitive writing.

These are not universal trade-offs; good systems can improve several dimensions together. The warning is to inspect what the chosen metric makes invisible.

Before setting a target, ask what behaviour could improve this number while making the actual objective worse. That adversarial question often reveals the missing guardrail.

Adding more metrics can create a new game

A common repair is to create a dashboard with ten targets. This can help if the measures represent genuinely important dimensions. It can also create a complicated scoring game where users optimise the formula rather than the work.

Keep a small set of complementary measures and retain direct review of the underlying outcome. Metrics should triangulate reality, not bury it under arithmetic.

Where possible, include measures that are difficult to improve without real progress: fresh transfer tasks, independent verification or user outcomes rather than only activity counts.

Targets can change the population being measured

Suppose a programme is rewarded for its pass rate. One route to improvement is better teaching. Another is accepting only students already likely to pass.

The reported percentage rises, but the population has changed. This connects to aggregation and composition effects: a metric can move because who enters the denominator changed.

Track access and entry criteria when the objective includes serving a broader population. Otherwise the system may improve the measured outcome by avoiding difficult cases.

Do not punish people for revealing problems

If a team is judged only by the number of reported errors, it may reduce reports instead of errors. A healthy system distinguishes discovering a defect from causing it.

In learning, a detailed error log can initially make performance look worse because more weaknesses become visible. That can be progress in diagnosis.

Rewarding absence of reported problems can therefore make the system blind. Pair outcome measures with mechanisms that encourage honest detection and correction.

Use metrics diagnostically, not ceremonially

A metric should trigger a question. If completion volume rises while transfer does not, ask what kind of questions are being selected. If average marks rise while harder tasks disappear, inspect the task mix.

Do not celebrate movement simply because the dashboard is green. Equally, do not dismiss a metric whenever it produces an inconvenient result. Investigate its relationship to the goal.

The value of measurement lies in better decisions, not in producing a number that can be displayed every week.

The repair routine

Write the actual objective without using the metric’s name. Then write what the metric measures directly. Identify at least one way the metric could improve while the objective stays unchanged or worsens.

Add a complementary check for that failure mode. Review changes in population, definitions and task difficulty. Keep direct samples of the underlying work so the metric can be audited against reality.

Finally, revisit the metric after people have adapted to it. A proxy that worked before incentives changed may need redesign once optimisation behaviour becomes visible.

For data interpretation, continue through the Mathematics Article Directory. For study systems, use Why More Practice Can Make the Wrong Method Harder to Fix.

Measure what helps you see the goal. The moment the number becomes more important than the thing it was meant to represent, the system has started working for the dashboard instead of the outcome.

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading