Why Comparing Marks From Different Tests Can Fail | Difficulty, Coverage and Real Progress

A higher percentage on a different test does not, by itself, prove that learning improved. A lower percentage does not, by itself, prove that learning declined. Before comparing marks, check what each paper assessed, how demanding the questions were, what help was available and how answers were marked. Then look for comparable pieces of work. The repair is not to stop using scores. It is to stop asking unlike assessments to behave like the same measuring instrument.

This guide is for the moment a student, parent or teacher places two papers side by side and asks, “Are we getting anywhere?” It explains how to make that conversation more accurate without turning a home revision session into a psychometrics project. All students, papers and numbers in the worked examples below are illustrative, not records of actual students or published assessment data.

Two papers arrive with one alarming story

Imagine Alicia bringing home a Mathematics paper marked 68 percent. The previous paper showed 80 percent. A familiar story forms immediately: the revision is not working, the student has become careless, or the extra lessons have made no difference.

Perhaps something really has deteriorated. But first inspect the papers. The earlier test covered two recently taught topics, allowed generous time and grouped similar questions together. The later paper mixed six topics and required students to choose methods without chapter headings. Some questions combined two ideas. Both papers were marked out of one hundred, but they did not ask the same questions of the learner.

The statement “the recorded percentage fell by twelve percentage points” is supported. The statement “Alicia knows twelve percent less Mathematics” is not. That second sentence treats a reporting number as a direct inventory of the knowledge inside a person.

This distinction is practical, not pedantic. The wrong interpretation can lead to the wrong intervention: more memorisation when method selection needs work, or an unnecessary programme change when the assessment itself became more demanding.

A percentage standardises the denominator, not the difficulty

Suppose one paper gives 18 marks out of 25 and another gives 36 out of 50. Both become 72 percent. Converting to percentages makes the fractions easier to compare arithmetically. It does not establish that the questions, scoring rules or required capabilities are equivalent.

Think of two walks, each described as ten kilometres. The common unit allows comparison of distance. It does not make their hills, surfaces or weather identical. Likewise, “out of one hundred” standardises how a score is expressed without standardising everything required to earn it.

Formal assessment programmes use evidence and statistical procedures when they need scores from different forms to be comparable. The ETS research memorandum Psychometric Evidence to Assess Expanded Test Use examines why score interpretations and comparisons require support, including conditions affecting equating across forms and populations. Its technical analysis is not a recipe for converting classroom percentages at home. It supports a narrower caution: a shared numerical scale is not sufficient evidence of equivalent meaning.

Do not invent a correction such as “add ten marks because that paper looked harder”. Unless a legitimate assessment provider supplies an appropriate conversion, retain the actual score and describe the differences in conditions separately.

The surprising case: both skills improve but the total falls

A change in topic weighting can reverse the story told by the total. Consider two hypothetical tests with routine questions and application questions. For simplicity, imagine that question quality and marking within each category are sufficiently comparable for this illustration.

On Test A, routine work carries 80 marks. The student earns 72 of them, or 90 percent. Application work carries 20 marks. The student earns 8, or 40 percent. The total is 80 out of 100.

On Test B, routine work carries only 20 marks. The student earns 19, or 95 percent. Application work now carries 80 marks. The student earns 40, or 50 percent. The total is 59 out of 100.

The total has fallen from 80 to 59, even though the recorded success rate rose in both categories. The second test placed much more weight on the weaker area. The arithmetic contains no contradiction: Test A rewards the stronger capability more heavily; Test B rewards the weaker capability more heavily.

This example does not prove that a real student improved whenever a total falls. Real papers rarely provide perfectly matched categories. It demonstrates why the composition of a score matters. Without checking composition, adults can mistake a changed assessment mixture for a changed learner.

The sensible response is twofold. Recognise the encouraging category evidence, while still addressing the substantial application weakness. Neither “everything is fine” nor “everything has collapsed” adequately describes the example.

Inspect what the question makes the student decide

Topic names can conceal important differences. Two questions may both be labelled percentages while demanding different decisions. One asks directly for ten percent of a stated quantity. Another asks for an original quantity after a percentage decrease. A third embeds the relationship in a comparison containing irrelevant information.

The learner may remember the relevant calculation but struggle to identify the starting quantity. A lower score on the later paper might therefore expose a selection problem rather than a failure to remember the basic operation.

In Science, naming a process and explaining an unfamiliar observation are not interchangeable tasks. In English, locating a stated fact and defending an inference from several sentences require different responses. In writing, a familiar narrative prompt and a tightly constrained communication task may expose different strengths.

When comparing questions, ask what the learner had to do before producing an answer. Did the question identify the method? Did it require combining information? Was the necessary evidence nearby? Were there plausible competing interpretations? These features tell you more than a chapter title alone.

The purpose is not to argue away every lost mark. It is to locate the demand that current practice has not yet prepared the student to meet.

Put the testing conditions beside the score

Support changes what a performance demonstrates. An answer produced with notes and a worked example can show useful learning in progress. It does not provide the same evidence as an answer produced later without those supports.

Record whether the attempt was independent, prompted or completed after seeing a solution. Also record whether the paper was interrupted, whether time was restricted and whether the student had previously attempted the questions. Keep this information factual and neutral. “Two prompts on question four” is more useful than “not really independent”.

Necessary access arrangements should not be casually removed to make a test look tougher. A learner may legitimately need an agreed format, assistive tool or accommodation to demonstrate the intended capability. Follow the relevant school or assessment arrangements. Document changes rather than treating support as automatically improper.

A fair comparison concerns the capability you intend to observe. If you are evaluating independent method selection, an adult naming the method changes that evidence. If you are evaluating reasoning with an authorised tool, using the tool may be entirely appropriate.

The question is not whether all help is bad. It is whether the conditions match the conclusion being drawn.

Check whether the marking changed

A writing percentage is not simply a count of correct facts. It can reflect several criteria, such as task fulfilment, organisation and language control. A new rubric, different weighting or different interpretation of a criterion may change the recorded score.

Before concluding that writing has declined, compare the actual work against the stated criteria. Did the student omit a required detail? Are sentences less controlled? Is evidence better connected to the argument? These are inspectable questions.

Where an important difference remains unclear, ask the teacher to explain the criterion using examples. Do not assume that the stricter-looking mark is wrong, and do not replace professional marking with a more flattering home score.

For self-marked practice, retain the marking guide and record uncertain decisions. A rise created by more generous self-marking is not the same as a rise in performance. Conversely, a newly accurate marking process may lower the reported score while improving the quality of the diagnosis.

Build a comparison record that can actually be used

A useful record can fit on one page. It needs the date, paper or task, topic coverage, conditions, score and a short note about recurring errors. Avoid an elaborate tracking system that consumes the time intended for learning.

First write the observation without interpretation: “24 out of 40 on a mixed paper; finished all questions; no notes.” Then identify one or two differences from the previous assessment: “Earlier paper was a single topic; current paper included graph interpretation.” Finally, state the question that remains: “Can the student interpret a scale correctly on a fresh graph?”

That final question creates a useful next action. A general verdict such as “weak in Science” does not. The record should reduce uncertainty about what to teach or practise, rather than merely preserve an emotional history of marks.

Where possible, keep a sample of the actual response. “Evidence omitted” becomes easier to understand when the student can see the sentence that made a claim without support. Preserve privacy when storing or sharing schoolwork; names and school identifiers are usually unnecessary for a learning review.

Compare a capability, not just two totals

After inspecting the papers, select a narrow capability that appears in both. It might be setting up a proportion, tracing a pronoun, explaining a controlled variable or checking a unit. The narrower question is easier to investigate responsibly.

Find several suitable examples, not just the single item that best supports the story you already prefer. Compare the demand, support and marking. If the items are too different, say so and use fresh tasks to clarify the current state instead.

Look at how the answer was produced. A correct answer with an unsupported guess is different evidence from a correct answer with sound reasoning. An incorrect final number can still contain an improved setup. Record both the improvement and the remaining error rather than forcing the work into one undifferentiated category.

For example, a learner may now choose the right equation but still make a transcription error. That is not full mastery, yet the next intervention should differ from the one used when the equation could not be formed at all.

The broader principle is explored in why chasing marks alone can hide the learning problem. Here the extra safeguard is to inspect whether the measurements themselves are comparable before interpreting their movement.

A small follow-up check is not formal test equating

A teacher can prepare a few fresh tasks that target the same skill and use a consistent marking approach. This can provide better local evidence than comparing unrelated totals. It still does not create a statistically equated examination scale.

Use similar demands without making the new questions cosmetic copies. Changing only the numbers can leave the student remembering the previous route. Changing everything at once can make the new task much harder. The aim is a reasonable instructional comparison, supported by professional judgement and transparent limitations.

Also consider the amount of evidence. One correct answer is encouraging but fragile. Several successful examples across different occasions are more informative about consistency. There is no universal number of questions that proves permanent mastery for every subject and learner.

Shared or repeated questions can provide clues, but prior exposure may influence later performance. Do not treat a remembered answer as an independent confirmation. Keep the repeat as practice and use an unseen task for the stronger check.

What to say when a parent asks, “Has this worked?”

A responsible answer separates the observed score, the evidence of capability and the uncertainty about cause. Consider this illustrative response:

“The overall percentage fell, but the second paper placed more marks on unfamiliar application. On the comparable algebra items, the student set up the relationships more consistently. Errors now appear mainly during rearrangement. We will check that on fresh mixed questions before changing the wider plan.”

This answer does not promise improvement that has not occurred. It also avoids letting one total erase relevant evidence. It describes what is known, what still fails and how the next observation will help.

Even a convincing improvement in capability does not, by itself, identify which intervention caused it. School teaching, independent practice, feedback and the passage of time may all be involved. “Performance improved while this approach was used” is often more defensible than “this one technique produced the entire gain”.

For families reviewing changes over time, the Parent Learning Support Directory provides wider routes without turning a single score into a diagnosis of the child.

Three quick cases to practise the distinction

Case one: a student scores 45 out of 50 on a chapter worksheet with notes and 35 out of 50 on an unfamiliar closed-book quiz. What can be said? The recorded score is lower under the second set of conditions. The difference does not isolate forgetting, because the task and support changed too. A fresh, appropriately matched independent task is a better next check.

Case two: a student earns 12 out of 20 on one response and 24 out of 40 on another. Both are 60 percent. Can we conclude that writing stayed unchanged? No. Check the criteria, task demands and actual writing. The equal percentages establish arithmetic equality, not equivalent quality.

Case three: a student improves from 6 to 9 correct answers on ten unfamiliar items designed to practise the same skill under similar conditions. Is that improvement? It is encouraging evidence on those tasks. It is not yet a guarantee about a whole examination, a permanent learning change or the cause of the gain.

These cases share a useful discipline: state the strongest conclusion the evidence supports, then stop before the claim becomes larger than the comparison.

What about rankings, averages and pass marks?

A class average can provide context, but it does not independently reveal how much one student learned. If everyone faces a more difficult paper, many scores may fall together. If class composition changes, a rank can change even when an individual produces similar work.

Do not mechanically subtract the class average from a score and label the result a precise measure of learning. That comparison describes relative performance in that group on that assessment. Its interpretation depends on the group, task and scoring system.

Likewise, a pass mark is a decision threshold. Crossing it matters where a school or examination uses that threshold, but one extra mark does not imply a sudden transformation in understanding. Being just below it does not make every underlying capability weak.

Keep the practical consequences and the educational diagnosis separate. A student may need to meet an official requirement while still benefiting from a much more detailed account of what should improve next.

Be careful when averaging percentages

Another comparison problem appears when a revision tracker averages percentages from tests of different lengths. Imagine two hypothetical results: 9 out of 10 and 25 out of 50. Their percentages are 90 and 50. The simple average of those percentages is 70. Combining the earned marks instead gives 34 out of 60, approximately 56.7 percent.

The calculations answer different questions. The first gives each assessment equal weight. The second gives each available mark equal weight. Neither automatically gives the correct school grade, and neither repairs differences in content or difficulty. Follow the actual assessment policy when calculating a formal result.

For learning diagnosis, preserve the separate observations rather than hiding them inside an unexplained average. Otherwise a short, easy quiz can offset a much longer, demanding paper in a way the reader never notices. A neat dashboard can therefore become less informative than two clearly labelled rows.

Also distinguish percentage points from percentage change. Moving from 50 percent to 60 percent is an increase of ten percentage points. Relative to the original score, the numerical increase is twenty percent. Neither expression means that the learner acquired twenty percent more of the entire subject. Report the score change plainly and explain the task before interpreting it.

Do not build the comparison around a convenient starting day

A very poor result can prompt a new study plan. The next result may be higher partly because the first attempt occurred under unusually difficult conditions. Conversely, an unusually strong first performance can make ordinary later work look like decline.

This does not make the new plan ineffective. It means the first observation should not carry more meaning than it can support. Where existing work is available, inspect several earlier attempts rather than choosing only the lowest score as the baseline for an impressive improvement story.

The same restraint applies when choosing the later result. Keep disappointing checks as well as successful ones. A learner who performs securely on one day and struggles on two others may need consistency work that would disappear from a record containing only the best performance.

Agree on the review question before inspecting the outcome. For instance, look at every suitable inference response from the selected practice period, not only the questions answered correctly. This simple habit makes the review fairer to the student and less vulnerable to the adult’s preferred explanation.

Turn the next fortnight into a better question

Begin with one high-value uncertainty from the marked work. For example: “Does the student still miss the relevant comparison when the wording changes?” Choose practice that makes that decision visible, give targeted feedback and preserve a few fresh questions for later checking.

At the next review, compare like with like as far as reasonably possible. Keep the task demand, support and marking sufficiently consistent to make the observation useful. Note unavoidable differences rather than quietly ignoring them.

Then make a bounded decision. Continue the practice if the targeted performance is becoming more reliable. Change the explanation or prerequisite work if the same misunderstanding persists. Seek the teacher’s view when marking or assessment expectations remain unclear. Avoid replacing the entire study programme because of an unexamined percentage.

This is a practical review sequence, not a validated two-week intervention or a promise of a particular mark increase. The calendar creates a review point; the evidence determines the judgement.

The number is useful when its meaning stays attached

Comparing test marks fails when the conditions, content and scoring disappear, leaving two percentages to carry the whole story. Restore those details and the same papers become much more informative.

Keep the actual marks. Inspect the assessed capabilities. Compare suitable work. Record support. Check again with fresh tasks. Make claims that fit the evidence.

A student deserves neither false reassurance from an easy paper nor a sweeping judgement from a difficult one. The useful question is not merely whether the number moved. It is what the student can now do, under what conditions, and what the next piece of work should help them do better.

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading