A combined percentage can reverse the comparison seen inside each subgroup. This can happen when the groups being combined have different baseline difficulty or success rates and the two options contain different proportions of those groups. Before trusting the total, break the data into relevant categories and inspect how the composition changed.
This pattern is often known as Simpson’s paradox. The name matters less than the habit: when an aggregate result seems surprising, ask what is inside the total before inventing an explanation.
A worked reversal
Imagine two fictional revision methods, A and B, used on easier and harder practice sets. On the easier sets, A records 9 successes out of 10 attempts, or 90 percent. B records 80 successes out of 100, or 80 percent. A has the higher observed success rate.
On the harder sets, A records 30 successes out of 100, or 30 percent. B records 2 successes out of 10, or 20 percent. Again, A has the higher observed success rate.
Now combine the groups. A has 39 successes out of 110 attempts, approximately 35.5 percent. B has 82 successes out of 110, approximately 74.5 percent. The aggregate comparison points in the opposite direction.
The arithmetic is correct. A handled mostly hard attempts; B handled mostly easy attempts. The total is heavily influenced by that composition.
The total answers a different question
The subgroup comparison asks how A and B performed within the easier category and within the harder category. The aggregate asks what fraction of all recorded attempts succeeded under each method as actually distributed across categories.
Those questions are not identical. The aggregate includes both method performance and the mixture of difficulty levels. If the mixture differs sharply, it can dominate the result.
Do not call the total “wrong”. It accurately describes the pooled observations. The failure occurs when it is interpreted as though the two methods faced comparable compositions and therefore isolates method quality.
Why simple averaging of subgroup percentages can also fail
For method A, the simple average of 90 percent and 30 percent is 60 percent. But its pooled success rate is about 35.5 percent because the hard category contains ten times as many A attempts as the easy category.
A simple average gives the two subgroup percentages equal weight. Pooling gives each attempt equal weight. Neither weighting rule is universally correct. The appropriate summary depends on the question and the population or scenario you intend to represent.
If a target population is half easy and half hard, an equal-weight standardised comparison may be relevant. If the task is to describe the actual observed workload, the pooled count may be relevant. State the weighting rather than letting it remain invisible.
Look for a variable connected to both grouping and outcome
Difficulty matters because it affects the chance of success and because the methods were used in very different proportions across difficulty levels. That makes it a natural variable to inspect.
In other settings, the relevant category could be age, prior attainment, task type, time period, location or another factor. Do not split data into every possible category until a preferred result appears. Choose categories because subject knowledge and the data-generating process make them relevant.
A subgroup variable can also be affected by the treatment or decision itself, in which case conditioning on it may create another problem. Aggregation repair is not a mechanical instruction to control for everything.
This pattern does not automatically prove causation
In the fictional example, we cannot conclude that method A causes higher success within either difficulty group. The attempts may differ in other ways, and the assignment of methods may not be random.
The subgroup table repairs a descriptive comparison by making difficulty composition visible. Causal claims require stronger assumptions or design.
This distinction prevents a common overcorrection: discovering Simpson’s paradox does not reveal which causal story is true. It reveals that the aggregate alone was insufficient for the interpretation being attempted.
A school example with subject mix
Suppose two study programmes are compared using Mathematics and English practice. Programme A is used mostly for difficult Mathematics sets; programme B mostly for easier English sets. A pooled score may favour B even if A has higher observed success within each subject.
Before claiming that B is the better programme, compare like subject with like subject and inspect difficulty within those subjects. The correct categorisation depends on how the tasks were assigned.
The reverse caution also applies. If the real-world programme will genuinely face a different mixture of tasks, its aggregate operational performance can matter. Do not discard the total simply because a stratified table exists.
A time trend can reverse for the same reason
Imagine a school increasingly gives a challenging assessment to a larger proportion of students over time. The overall pass percentage may fall even if pass rates improve within both the easier and harder assessment categories.
The changing mixture can create the aggregate decline. Conversely, an overall rise can occur when the population shifts towards an easier category even if subgroup performance is unchanged or worse.
When comparing years, therefore, inspect whether the tested population, task mix or definitions changed. The guide to comparing marks from different tests addresses a closely related assessment problem.
Graphs can hide the mixture
A single pair of bars showing 35.5 percent versus 74.5 percent makes B look overwhelmingly stronger in the fictional example. Adding the easy and hard subgroup bars reveals that the comparison changes inside each category.
The correct response is not to choose whichever graph supports the desired story. Show the structure relevant to the question. Where the aggregate and subgroup views differ, explain why.
Continue with Why Comparing Graphs by Eye Can Fail for axis, scale and denominator problems that can compound the aggregation issue.
Standardisation makes the weighting decision explicit
If the purpose is to compare methods under the same hypothetical mix of easy and hard tasks, choose a common set of weights. For example, an equal 50–50 mix gives A an observed weighted rate of 0.5×90% + 0.5×30% = 60%. B gives 0.5×80% + 0.5×20% = 50%.
Those standardised values answer a constructed question: what would the rates be under this chosen mixture, using the observed subgroup rates? They are not the actual pooled rates and do not remove uncertainty or confounding within the subgroups.
Use a target mixture justified by the decision, not one selected because it produces the preferred ranking. Report the weights so another reader can reproduce the comparison.
Small subgroups deserve caution
In our invented table, some subgroups contain only ten attempts. Their percentages can move substantially when one outcome changes. The example is designed to make the reversal easy to see, not to provide a statistically precise estimate of method performance.
Real decisions should consider sample size, uncertainty, repeated observations and study design. A dramatic subgroup percentage from a tiny cell should not be treated as a stable law merely because it repairs an aggregate comparison.
The lesson is structural: inspect composition. The amount of confidence appropriate for the resulting rates is a further question.
The repair routine
First calculate the aggregate correctly. Then identify plausible categories that affect both the outcome and the composition of the compared groups. Inspect the counts and rates within those categories.
Ask which question the decision actually needs: observed total performance, within-category comparison, or performance under a common target mixture. Keep all three distinct if more than one matters.
Do not average percentages without stating their weights. Do not split the data after the fact solely to manufacture a preferred conclusion. Do not interpret a descriptive reversal as automatic proof of a causal mechanism.
For the denominator problem underneath many aggregate errors, use Why Comparing Rates With Different Denominators Can Mislead. For broader statistical reasoning, continue through the Mathematics Article Directory.
A total is powerful because it compresses information. That same compression can hide the structure that changes its meaning. When the conclusion feels surprising, open the total and look inside.
