How Inter-Rater Reliability and Agreement Work | When Two People Judge the Same Evidence

Two teachers read the same essay. One awards 17 marks; the other awards 13. Two archivists inspect the same photograph. One labels it “industrial infrastructure”; the other chooses “transport infrastructure”. Two researchers code the same interview. One sees an example of trust; the other sees compliance.

The disagreement is not automatically evidence that one person is careless. It may reveal an ambiguous rubric, a category boundary that was never defined, a task that asks for more precision than the evidence can support, or a judgement that genuinely contains more than one defensible interpretation.

Inter-rater agreement asks how closely different raters give the same or sufficiently similar judgements. Inter-rater reliability asks how well a measurement process distinguishes the things being rated despite variation between raters. The two ideas overlap, but they are not interchangeable. A study can have high consistency without exact agreement, and high raw agreement can coexist with a weak chance-corrected reliability coefficient.

This guide explains the full judgement chain: object → rubric → rater interpretation → recorded judgement → comparison → reliability or agreement statistic → diagnosis → repair. The aim is not to find one universal coefficient. It is to make human judgement measurable enough that readers can see what was judged, what counted as agreement and which source of variation remains unresolved.

Reading route: begin with agreement versus reliability, work through the kappa example and its prevalence problem, then move to continuous ratings and ICCs, study design, rater drift and the reporting contract.

Agreement and reliability answer different questions

Suppose two judges score essays on a 0–20 scale. Judge B consistently gives every essay two marks more than Judge A. The rank order of essays may be almost identical: the strongest essay for A is the strongest for B, and so on. That is strong consistency. It is not exact agreement.

If the score will be used to rank essays, consistency may be important. If the score determines whether a learner crosses a fixed threshold, a two-mark systematic difference can matter greatly. The correct statistic therefore depends on the use of the judgement.

Agreement is often expressed in the units of the decision: same category, same score, within one mark, within five millimetres, or within a predeclared tolerance. Reliability is usually a ratio or coefficient describing how much observed variation reflects real differences between rated objects rather than rater-related or measurement noise.

The first editorial discipline is to state which of these jobs matters. “We checked reliability” is too vague when the decision actually requires near-exact scoring agreement.

A rating is a measurement process, not merely an opinion

A rater does not act on the object alone. The rating is produced by an interaction among the object, instructions, definitions, examples, training, time pressure, interface and rater. Change any of those and the observed judgement may change.

That means an inter-rater study is partly a test of the instrument. If ten intelligent raters disagree because the category “advanced” was never operationally defined, the problem may sit in the rubric rather than in the people.

Conversely, a beautifully written rubric can still be used inconsistently if raters receive different examples or interpret exception rules differently. Reliability belongs to the complete measurement arrangement under specified conditions.

This connects directly to Construct Validity and Measurement Models: before asking whether raters agree, establish what their score or category is intended to represent.

Percentage agreement is useful—and incomplete

For categorical ratings, the simplest measure is often the percentage of items on which raters give the same category. It is direct, interpretable and should not be discarded merely because more sophisticated statistics exist.

If two coders independently classify 100 records and agree on 85, observed agreement is 85%. That tells a reader something concrete. It does not tell us how much agreement would be expected merely from the raters’ marginal category frequencies, nor whether disagreement is concentrated in one category.

Percentage agreement is therefore best treated as a visible first layer. Keep the confusion table, not only the percentage. A single summary can hide whether one rater systematically overuses a category or whether nearly all disagreements concern one difficult boundary.

A worked kappa example: 85% agreement is not the whole story

Imagine two raters classify 100 items as pass or revise. Their results are:

Constructed two-rater classification example
Rater B: PassRater B: ReviseTotal
Rater A: Pass65570
Rater A: Revise102030
Total7525100

Observed agreement is (65 + 20) ÷ 100 = 0.85. If the raters independently maintained those same marginal proportions, expected agreement would be (0.70 × 0.75) + (0.30 × 0.25) = 0.60.

Cohen’s kappa is (observed agreement − expected agreement) ÷ (1 − expected agreement). Here that is (0.85 − 0.60) ÷ 0.40 = 0.625.

The original Cohen paper on agreement for nominal scales introduced this chance-corrected logic. The coefficient is not a percentage of correct decisions and should not be described as such. It measures agreement relative to the agreement expected from the raters’ marginal classifications under the statistic’s model.

Why high raw agreement can coexist with low kappa

Now imagine that almost every item belongs to one category. Both raters label 95 of 100 items positive. They agree that 90 are positive, and each assigns five of the remaining items differently. They agree on 90% of all items.

Because each rater uses the positive category 95% of the time, the expected agreement from the marginals is 0.95² + 0.05² = 0.905. Kappa becomes (0.90 − 0.905) ÷ (1 − 0.905), about −0.053.

This does not prove that the raters are wildly incompetent. It says that their 90% raw agreement is no better than the agreement implied by their extremely imbalanced marginal classifications under the kappa model. The rare category is precisely where they fail to agree.

That distinction can be substantively important. If the rare category marks the cases that require intervention, 90% overall agreement may be operationally poor. If the rare category is inconsequential, the same table may have a different practical meaning.

Do not respond to a surprising kappa by automatically replacing it with raw agreement or by declaring the statistic defective. Inspect the table, category prevalence, rater marginals and decision consequences. The disagreement pattern contains more information than the coefficient alone.

Nominal, ordinal and continuous ratings need different agreement logic

Categories have structure. Misclassifying excellent as good may be less serious than misclassifying excellent as inadequate. A nominal kappa treats every disagreement as simply disagreement. Weighted kappa allows an ordinal rating system to assign different penalties to different distances between categories.

The weights are part of the analysis. Linear and quadratic weighting embody different assumptions about how disagreement grows with category distance. They should not be selected after seeing which one produces a more attractive coefficient.

For continuous ratings, collapsing scores into categories merely to calculate kappa throws away information. Intraclass correlation coefficients, limits of agreement and direct error summaries are often more suitable, depending on the measurement job.

Correlation is not agreement

Two raters can have a Pearson correlation of 1 while disagreeing on every absolute score. If B = A + 2 for every item, their linear correlation is perfect because every change in A is mirrored by B. Yet every score differs by two.

Correlation measures association. It does not, by itself, test whether two measurements are interchangeable. This is why method reports that present only a correlation between raters can give false reassurance when absolute agreement matters.

A strong analysis chooses a statistic that matches the intended use: ranking, consistency, exact score agreement or agreement within a tolerance.

The ICC is a family, not one number

Intraclass correlation coefficients are widely used for continuous or near-continuous ratings, but “the ICC” is not a single formula. Different ICC forms correspond to different study designs and inferential targets.

The practical questions include: were all items rated by the same raters? Are those raters the only raters of interest, or are they viewed as representatives of a wider rater population? Does the decision require absolute agreement or only consistency? Will the final measurement use one rater or the average of several raters?

Koo and Li’s guideline on selecting and reporting ICCs emphasises that model, type and definition should be reported because different choices lead to different interpretations. Their article also has a published erratum, a useful reminder that even methodological references require version awareness.

The 2026-accessed COSMIN manual similarly distinguishes agreement-oriented and consistency-oriented ICC choices in its measurement-property setting. COSMIN is designed for health measurement instruments; its statistical distinctions are useful here as methodological examples, not as a reason to turn a general educational rating problem into a clinical one.

Absolute agreement and consistency can lead to different ICC conclusions

Return to two essay markers. If one marker is always two points higher, a consistency-oriented coefficient can remain high because the relative ordering is preserved. An absolute-agreement coefficient treats the systematic offset as disagreement.

Neither is universally superior. A research team studying whether raters rank specimens similarly may care about consistency. An examination system using a fixed pass mark usually cares about absolute score differences as well.

The statistical model should therefore be selected from the decision purpose, not the other way around. “Our ICC was 0.90” is incomplete unless the reader knows which ICC and why it fits the rating design.

Single-rater and average-rater reliability are different objects

A mean of three independent ratings can be more stable than one rating because some rater-specific variation averages out. An ICC for the average of several raters can therefore be higher than the corresponding single-rater ICC.

That higher coefficient does not show that any individual rater has become more reliable. It describes the reliability of a different measurement procedure: the average of several ratings.

Report the procedure that will actually be used. A study should not validate a three-rater average and then deploy one unaided rater while quoting the larger reliability coefficient.

Reliability depends on variation among the rated objects

A reliability coefficient is partly a ratio of between-object variation to total variation. If every item is extremely similar, even modest rater noise can occupy a large share of the observed variance. The reliability coefficient may therefore be lower than in a more heterogeneous sample, even when absolute rater error is unchanged.

This is why reliability is not solely an intrinsic property of a rater. It belongs to a measurement procedure applied to a population of objects under specified conditions.

A study using only obvious examples can also create the opposite problem: excellent agreement on easy cases may not represent the difficult boundary cases for which the rubric is actually needed.

Agreement within a tolerance may be the real operational question

For continuous measurements, exact equality may be unnecessarily strict. Two measurements of 14.01 and 14.02 can be practically interchangeable even though they are not identical. In other settings, a one-unit difference can change a classification and be unacceptable.

Define a tolerance from the use case. Then report the proportion of ratings within that tolerance, the distribution of differences and whether disagreement is systematically biased in one direction. If a tolerance is chosen after inspecting the data, say so; do not present it as a pre-existing standard.

NIST’s work on interlaboratory comparisons illustrates a broader measurement principle: systematic differences and random variation among measurement sources are distinct components worth diagnosing separately.

Rater training changes the instrument

Training is not an incidental administrative step. Examples, counterexamples, anchor responses and discussion of boundary cases alter how raters interpret the rubric. A reliability study conducted before calibration describes a different procedure from one conducted afterward.

Good training should make distinctions clearer without turning judgement into imitation. If raters are shown the final answer to every study item before the reliability test, agreement becomes unsurprising and uninformative.

Separate training material from evaluation material where possible. Record which rules were revised after calibration. A rubric that becomes clearer through disagreement has improved; the first study has still served a useful diagnostic purpose.

Independence matters during the reliability assessment

If raters discuss an item before recording their individual judgements, the resulting consensus does not measure independent inter-rater agreement. It measures a joint decision process.

Both can be valuable. A production workflow may deliberately use discussion to resolve difficult cases. But the reliability study should distinguish independent first ratings from adjudicated final ratings.

Adjudication can improve the final record while hiding where the rubric generates disagreement. Preserve both stages when the purpose includes instrument improvement.

Adjudication is not a substitute for reliability evidence

A common workflow lets two raters score independently and sends disagreements to a senior adjudicator. The final dataset then contains one clean answer per item. If only those final answers are retained, later users may wrongly infer that the original coding was unambiguous.

Keep a disagreement flag, original ratings and adjudication rule where appropriate. The rate and pattern of adjudication can itself be a useful quality signal.

A senior adjudicator is also a rater. Their decision may be authoritative because of governance, expertise or a defined hierarchy, but it is not automatically ground truth.

A reference standard changes the question from agreement to accuracy

If an independent, defensible reference label exists, compare each rater with that reference to study accuracy or classification performance. Two raters can agree perfectly and still both be wrong.

Conversely, two raters can disagree because one is closer to a valid reference standard. An agreement coefficient cannot tell us which one.

Do not call the senior rater a gold standard merely because that person is senior. A reference standard needs its own evidential basis.

Design the reliability study around the deployment conditions

The sample should include the kinds of objects the rating system will actually encounter: easy, difficult, common and consequential boundary cases. If raters will work across several age groups, document whether the reliability study covers them. If the rubric will be used on new genres, do not infer reliability from one narrow genre without qualification.

Use enough items and raters to estimate the intended coefficient with useful precision. A reliability point estimate without an uncertainty interval can look more stable than the study supports. Sample-size planning depends on the coefficient, expected reliability, desired precision and design; there is no universal magic number of rated items.

Random or representative sampling of items can support broader inference, but deliberately sampling difficult cases can be valuable for stress testing. Label the purpose. A stress test is not a prevalence estimate.

Not every item needs every rater—but the design must be explicit

Large projects sometimes use incomplete rating designs: each item is scored by a subset of raters. This can be efficient, but it changes the statistical problem. The overlap structure must be sufficient to separate item and rater variation under the chosen model.

A disconnected design in which one set of raters sees only one set of items and another set sees entirely different items can make rater severity inseparable from item difficulty. A model may fit, but the comparison lacks a bridge.

Plan the assignment graph before data collection. Connectivity is a design property, not something software can restore afterward.

Missing ratings need their own explanation

A rater may skip an item because it is genuinely unrateable, because time expired, because the interface failed or because the item was especially difficult. These reasons are not equivalent.

Silently analysing only complete rating pairs can change the difficulty distribution. If the hardest items are also the most likely to be skipped, reliability among completed items may overstate reliability in the intended workflow.

Record missing-rating reasons where feasible and route the analysis through Missing Data Analysis when the pattern can alter inference.

Rater drift turns reliability into a time-dependent property

Raters learn, forget, reinterpret and adapt. A team can begin highly calibrated and gradually diverge. New examples can shift the perceived meaning of a category. Production pressure can encourage shortcuts.

Long-running systems therefore need periodic checks using fresh or repeated anchor items, with care to avoid simple memorisation. Plot disagreement over time and inspect category-specific changes rather than relying only on one annual coefficient.

If the rubric itself is revised, do not merge before-and-after reliability as though the measurement system were unchanged. A new edition deserves its own evidence.

More agreement is not always more validity

A team can become extremely consistent by adopting a narrow shortcut. If every rater learns to infer essay quality from length, agreement may rise while construct validity falls.

This is why reliability is necessary for many measurement uses but insufficient for validity. A perfectly reproducible wrong rule remains wrong.

Pair reliability evidence with content, structural, criterion or consequential evidence appropriate to the construct. The reliability coefficient should never be asked to prove what it was not designed to prove.

Machine-assisted rating does not remove the rater problem

An automated system can act as another measurement process. Comparing a model with human raters can be useful, but human consensus is not automatically ground truth. If humans share the same rubric ambiguity, a model trained on their labels may reproduce that ambiguity efficiently.

Evaluate machine-human agreement, human-human agreement and, where available, performance against an independent reference. Check whether disagreements concentrate in particular groups, formats or edge cases.

When a model changes version, rerun relevant reliability and validity checks. A software update can change the effective rater even when the user interface looks identical.

Reliability can be decomposed into sources of variation

Simple inter-rater coefficients collapse a complex process into one summary. More elaborate designs can separate object, rater, occasion and interaction components. Generalizability theory extends this idea by asking how multiple facets of measurement contribute to error for a defined decision.

The practical value is diagnostic. If rater severity dominates, calibration may help. If object-by-rater interaction dominates, examples or rubric boundaries may need work. If occasion variation dominates, fatigue or changing conditions may matter.

Do not add complexity merely to produce a more impressive model. Use decomposition when it changes what can be repaired or how the operational score should be formed.

A disagreement map is often more useful than one coefficient

For categorical tasks, show the full cross-tabulation. For ordinal tasks, inspect how far apart disagreements are. For continuous scores, plot one rater against another and inspect the distribution of differences.

Stratify carefully by meaningful features such as item type, language demand or difficulty. A global coefficient can hide a weak subgroup. But do not generate dozens of subgroup coefficients without a plan and then highlight only the surprising one; that creates a multiple-testing and interpretation problem of its own.

Use Data Visualisation to make the disagreement structure readable without turning uncertainty into decoration.

What to do when agreement is poor

Do not begin by removing the most disagreeable rater. Diagnose the mechanism. Are definitions unclear? Are categories overlapping? Are examples unrepresentative? Is the task asking raters to infer information not present in the object? Are one or two raters using a systematically different threshold?

Repair the smallest load-bearing defect. Rewrite a boundary rule, add counterexamples, change the scoring scale, separate two concepts, improve the interface or create an escalation route for genuinely unrateable items.

Then test the revised process on new items. Reusing only the cases discussed during training can measure memory rather than generalisable reliability.

A reporting contract for inter-rater studies

A reader should be able to reconstruct the rating system. Report the objects, sampling method, number and background of raters, training, independence, rubric edition, rating scale, missing-rating rules and whether raters were blinded to other information where that matters.

Report the raw agreement structure alongside the selected reliability statistic. For kappa, identify whether it is unweighted or weighted and state the weighting scheme. For ICCs, report the model, single-versus-average type and agreement-versus-consistency definition, following the design logic described in the ICC reporting guideline.

Include uncertainty intervals where appropriate. State what level of agreement was required for the intended use and where that threshold came from. A universal adjective such as “excellent” is less informative than a coefficient interpreted against a real decision.

Finally, preserve adjudication and calibration history separately from independent first ratings. That allows future users to distinguish measurement quality from the quality-control process built around it.

A learner’s six-question reliability check

  1. What exactly are the raters judging?
  2. What counts as agreement for the decision?
  3. Were the ratings independent before adjudication?
  4. Does the statistic match nominal, ordinal or continuous data?
  5. Could prevalence, sample composition or rater drift change the coefficient?
  6. Would high agreement still mean the judgement is valid?

These questions prevent a coefficient from becoming a ritual. The purpose of reliability analysis is not to decorate a methods section. It is to discover whether a human judgement process is stable enough for the responsibility being placed on it.

Where this article sits in the eduKate Library

This article owns the general methodological problem of comparing human raters. It does not replace subject-specific scoring standards, examination procedures, clinical measurement guidance or professional judgement frameworks. Those remain with their relevant canonical owners.

Use Measurement Error and Misclassification when the recorded value differs from the underlying thing, Construct Validity when the meaning of the score is the problem, and Statistical Inference and Uncertainty for interval estimation and inferential reasoning.

Sources and further reading

Source pages and accessible methodological records were checked for this edition on 5 September 2026. The worked 100-item tables are original examples. This article is an explanatory synthesis, not a validated scoring protocol for any specific examination or professional assessment.

Continue through the Library: Construct Validity and Measurement ModelsMeasurement Error and MisclassificationData VisualisationResearch Collections Directory.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading