Why Combining Groups Can Reverse a Comparison | When the Total Hides What Happened Inside

A combined percentage can reverse the comparison seen inside each subgroup. This can happen when the groups being combined have different baseline difficulty or success rates and the two options contain different proportions of those groups. Before trusting the total, break the data into relevant categories and inspect how the composition changed.

This pattern is often known as Simpson’s paradox. The name matters less than the habit: when an aggregate result seems surprising, ask what is inside the total before inventing an explanation.

A worked reversal

Imagine two fictional revision methods, A and B, used on easier and harder practice sets. On the easier sets, A records 9 successes out of 10 attempts, or 90 percent. B records 80 successes out of 100, or 80 percent. A has the higher observed success rate.

On the harder sets, A records 30 successes out of 100, or 30 percent. B records 2 successes out of 10, or 20 percent. Again, A has the higher observed success rate.

Now combine the groups. A has 39 successes out of 110 attempts, approximately 35.5 percent. B has 82 successes out of 110, approximately 74.5 percent. The aggregate comparison points in the opposite direction.

The arithmetic is correct. A handled mostly hard attempts; B handled mostly easy attempts. The total is heavily influenced by that composition.

The total answers a different question

The subgroup comparison asks how A and B performed within the easier category and within the harder category. The aggregate asks what fraction of all recorded attempts succeeded under each method as actually distributed across categories.

Those questions are not identical. The aggregate includes both method performance and the mixture of difficulty levels. If the mixture differs sharply, it can dominate the result.

Do not call the total “wrong”. It accurately describes the pooled observations. The failure occurs when it is interpreted as though the two methods faced comparable compositions and therefore isolates method quality.

Why simple averaging of subgroup percentages can also fail

For method A, the simple average of 90 percent and 30 percent is 60 percent. But its pooled success rate is about 35.5 percent because the hard category contains ten times as many A attempts as the easy category.

A simple average gives the two subgroup percentages equal weight. Pooling gives each attempt equal weight. Neither weighting rule is universally correct. The appropriate summary depends on the question and the population or scenario you intend to represent.

If a target population is half easy and half hard, an equal-weight standardised comparison may be relevant. If the task is to describe the actual observed workload, the pooled count may be relevant. State the weighting rather than letting it remain invisible.

Look for a variable connected to both grouping and outcome

Difficulty matters because it affects the chance of success and because the methods were used in very different proportions across difficulty levels. That makes it a natural variable to inspect.

In other settings, the relevant category could be age, prior attainment, task type, time period, location or another factor. Do not split data into every possible category until a preferred result appears. Choose categories because subject knowledge and the data-generating process make them relevant.

A subgroup variable can also be affected by the treatment or decision itself, in which case conditioning on it may create another problem. Aggregation repair is not a mechanical instruction to control for everything.

This pattern does not automatically prove causation

In the fictional example, we cannot conclude that method A causes higher success within either difficulty group. The attempts may differ in other ways, and the assignment of methods may not be random.

The subgroup table repairs a descriptive comparison by making difficulty composition visible. Causal claims require stronger assumptions or design.

This distinction prevents a common overcorrection: discovering Simpson’s paradox does not reveal which causal story is true. It reveals that the aggregate alone was insufficient for the interpretation being attempted.

A school example with subject mix

Suppose two study programmes are compared using Mathematics and English practice. Programme A is used mostly for difficult Mathematics sets; programme B mostly for easier English sets. A pooled score may favour B even if A has higher observed success within each subject.

Before claiming that B is the better programme, compare like subject with like subject and inspect difficulty within those subjects. The correct categorisation depends on how the tasks were assigned.

The reverse caution also applies. If the real-world programme will genuinely face a different mixture of tasks, its aggregate operational performance can matter. Do not discard the total simply because a stratified table exists.

A time trend can reverse for the same reason

Imagine a school increasingly gives a challenging assessment to a larger proportion of students over time. The overall pass percentage may fall even if pass rates improve within both the easier and harder assessment categories.

The changing mixture can create the aggregate decline. Conversely, an overall rise can occur when the population shifts towards an easier category even if subgroup performance is unchanged or worse.

When comparing years, therefore, inspect whether the tested population, task mix or definitions changed. The guide to comparing marks from different tests addresses a closely related assessment problem.

Graphs can hide the mixture

A single pair of bars showing 35.5 percent versus 74.5 percent makes B look overwhelmingly stronger in the fictional example. Adding the easy and hard subgroup bars reveals that the comparison changes inside each category.

The correct response is not to choose whichever graph supports the desired story. Show the structure relevant to the question. Where the aggregate and subgroup views differ, explain why.

Continue with Why Comparing Graphs by Eye Can Fail for axis, scale and denominator problems that can compound the aggregation issue.

Standardisation makes the weighting decision explicit

If the purpose is to compare methods under the same hypothetical mix of easy and hard tasks, choose a common set of weights. For example, an equal 50–50 mix gives A an observed weighted rate of 0.5×90% + 0.5×30% = 60%. B gives 0.5×80% + 0.5×20% = 50%.

Those standardised values answer a constructed question: what would the rates be under this chosen mixture, using the observed subgroup rates? They are not the actual pooled rates and do not remove uncertainty or confounding within the subgroups.

Use a target mixture justified by the decision, not one selected because it produces the preferred ranking. Report the weights so another reader can reproduce the comparison.

Small subgroups deserve caution

In our invented table, some subgroups contain only ten attempts. Their percentages can move substantially when one outcome changes. The example is designed to make the reversal easy to see, not to provide a statistically precise estimate of method performance.

Real decisions should consider sample size, uncertainty, repeated observations and study design. A dramatic subgroup percentage from a tiny cell should not be treated as a stable law merely because it repairs an aggregate comparison.

The lesson is structural: inspect composition. The amount of confidence appropriate for the resulting rates is a further question.

The repair routine

First calculate the aggregate correctly. Then identify plausible categories that affect both the outcome and the composition of the compared groups. Inspect the counts and rates within those categories.

Ask which question the decision actually needs: observed total performance, within-category comparison, or performance under a common target mixture. Keep all three distinct if more than one matters.

Do not average percentages without stating their weights. Do not split the data after the fact solely to manufacture a preferred conclusion. Do not interpret a descriptive reversal as automatic proof of a causal mechanism.

For the denominator problem underneath many aggregate errors, use Why Comparing Rates With Different Denominators Can Mislead. For broader statistical reasoning, continue through the Mathematics Article Directory.

A total is powerful because it compresses information. That same compression can hide the structure that changes its meaning. When the conclusion feels surprising, open the total and look inside.

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading