Two students can obtain the same raw score and still have answered very different questions correctly.
A test can also become harder even when the number of questions stays the same. A student can improve while a raw percentage remains unchanged because the item set changed. And two items with the same percentage-correct rate can carry different information about where a learner sits on an underlying proficiency scale.
Item response theory, or IRT, exists because measurement cannot always be reduced to counting correct answers. It models the relationship between a person’s position on a latent trait and the probability of particular responses to individual items. Rasch measurement is one influential member of that wider family.
Quick Read
- IRT models responses at the item level rather than treating every item as interchangeable.
- A latent trait such as proficiency is usually represented by a parameter often written as θ.
- Item difficulty describes where on the trait scale an item is most informative or challenging.
- Two- and three-parameter models can also represent discrimination and, for some item formats, a lower asymptote.
- Rasch models impose stronger constraints: item discrimination is fixed rather than freely estimated.
- Item information describes how precisely an item measures different parts of the latent scale.
- Differential item functioning matters because an item may behave differently across groups even after conditioning on the latent trait.
- IRT can support equating, adaptive testing and large-scale assessment, but only when model assumptions and linking conditions are defensible.
One-Sentence Answer
Item response theory works by modeling the probability of each item response as a function of an unobserved trait and item properties, allowing measurement precision, item difficulty, scale linking and score interpretation to vary across the trait rather than assuming all questions contribute equally.
1. Why Raw Scores Are Not Always Enough
Suppose a ten-item test contains five easy items and five hard items. Student A answers the five easy items correctly. Student B answers the five hard items correctly. Both obtain 5/10.
A raw-score system treats those response patterns as identical totals. An item-level model asks whether the pattern itself carries information about proficiency and about the items. That is the starting intuition behind IRT.
This does not mean raw scores are useless. In many well-designed tests they are transparent, reliable and highly correlated with model-based scores. IRT becomes especially valuable when we need to place different item sets on a common scale, understand item behaviour, target measurement precision or support adaptive administration.
2. Classical Test Theory and IRT Ask Different Questions
Classical test theory often begins with an observed score decomposed conceptually into a true-score component plus error. Reliability is usually treated at the test-score level, and item statistics such as difficulty or discrimination are often sample-dependent summaries.
IRT moves the model to the item-response level. Instead of asking only how reliable the total score is, it asks how the probability of a response changes with a latent trait and how much information each item provides at different trait levels.
The two frameworks are not enemies. They are different measurement lenses. A strong assessment program may use both.
3. The Latent Trait: What Is θ?
IRT commonly represents a person’s position on an underlying scale by the Greek letter theta, θ. In educational testing, θ may represent proficiency in a defined domain. In other settings it could represent an attitude, symptom severity or another latent continuum.
The latent trait is not observed directly. It is inferred from the pattern of item responses under the model. Its meaning therefore depends on the assessment framework, item content, model, population and scale construction.
A θ estimate should never be treated as a pure measurement of an invisible essence. It is a model-based position on a scale defined through evidence.
4. The Item Characteristic Curve
For a dichotomous item scored incorrect/correct, the item characteristic curve shows how the probability of a correct response changes as θ changes.
Easy items shift the curve so that lower levels of θ correspond to substantial success probabilities. Hard items shift it toward higher θ. Steeper curves discriminate more sharply between nearby levels of the trait around their transition region.
This curve is the visual heart of IRT. It turns an item from “worth one mark” into a measurement instrument whose behaviour varies across the proficiency scale.
5. The One-Parameter Logistic Model and Rasch Family
In a one-parameter logistic model, item difficulty varies but discrimination is constrained to be the same across items. The Rasch model is closely related and has a distinctive measurement philosophy: the data are expected to conform to a model with invariant item discrimination rather than allowing each item to acquire its own slope merely to improve fit.
That restriction is scientifically important. It creates stronger measurement requirements and, when supported, can aid comparability. But stronger constraints also mean more ways for real data to fail the model.
Rasch analysis should therefore not be reduced to “an easier IRT model”. It is a specific modeling and measurement stance about how item difficulty and person location relate.
6. The Two-Parameter Logistic Model
The two-parameter logistic, or 2PL, model allows each item to have its own discrimination parameter as well as its own difficulty parameter.
NAEP technical documentation describes the 2PL for dichotomously scored items with a slope parameter characterising sensitivity to the scale score and a threshold parameter characterising item difficulty.
An item with high discrimination has a steep transition: small differences in θ around the item’s difficulty produce larger differences in response probability. An item with low discrimination changes more gradually and therefore separates nearby proficiency levels less strongly.
7. The Three-Parameter Logistic Model
The 3PL adds a lower-asymptote parameter, often used for multiple-choice items where very low-proficiency respondents may still have a non-zero probability of answering correctly.
NAEP technical documentation describes its 3PL model in terms of discrimination, difficulty and a lower asymptote reflecting the chance of selecting the correct option at very low scale levels.
The parameter is sometimes casually called “guessing”, but that label can be misleading. The lower asymptote is a model parameter. Real response behaviour may reflect guessing, partial knowledge, distractor structure, test-taking strategies or other mechanisms.
8. Polytomous Items and the Generalised Partial Credit Model
Not all assessment responses are simply right or wrong. Constructed responses may receive 0, 1, 2, 3 or more score categories. Attitude scales may have ordered response options. Such items require models for multiple categories.
The generalized partial credit model, or GPCM, is one widely used IRT model for ordered score categories. NAEP uses polytomous IRT approaches for constructed-response items with multiple score points.
The key conceptual shift is that an item no longer has one transition from incorrect to correct. It has several category boundaries or step structures describing how response probabilities change across θ.
9. Item Difficulty Is Relative to the Scale
An item is not “hard” in an absolute universal sense. Difficulty is defined relative to a population, domain, scale and model. An algebra item may be difficult for one age group and easy for another. A translated item may change difficulty because language demand changes. An item can also drift over time as curricula, culture or exposure change.
IRT therefore treats item calibration as empirical measurement, not as a permanent property written into the question itself.
10. Item Information: Precision Is Not Constant
One of IRT’s most useful ideas is item information. An item provides more information where its response probability is most sensitive to changes in θ and less information where almost everyone answers correctly or almost everyone answers incorrectly.
Information accumulates across items. A test information function shows where the whole test measures precisely and where it measures poorly. Standard errors of θ are inversely related to the square root of information: more information means greater precision.
This is a major difference from the idea that one reliability coefficient describes measurement precision equally for everyone. An IRT scale can be highly precise around one proficiency range and weak at another.
11. Test Targeting
A well-targeted assessment contains enough item information where the intended population is expected to lie. If every item is easy, the test provides little precision among high-performing respondents. If every item is very difficult, it provides little precision at the lower end.
Targeting therefore connects psychometrics to assessment purpose. A screening tool, a mastery test and a competition exam should not necessarily concentrate information in the same part of the trait scale.
12. Local Independence
A common IRT assumption is that once the latent trait is conditioned on, responses to different items are independent, or sufficiently close to independent for the model’s purpose.
Local dependence can arise when several questions share the same reading passage, one question gives away the answer to another, items repeat nearly identical wording, or speed and fatigue create residual response patterns not captured by θ.
If local dependence is strong, the test may appear to contain more independent information than it really does. OECD’s work on PISA scaling explicitly discusses nuisance factors such as testlet effects and locally dependent items as issues requiring attention.
13. Dimensionality
Many standard IRT models assume that one dominant latent dimension explains the relevant response structure within a scale. Real assessments can involve multiple skills: vocabulary, reasoning, speed, background knowledge and strategy may all contribute.
“Unidimensional” should therefore be understood as a modeling claim, not a statement that human performance literally has only one cause. The question is whether a single latent dimension provides an adequate measurement representation for the intended score interpretation.
When it does not, multidimensional IRT or a different scale design may be more appropriate.
14. Calibration: Placing Items on a Common Scale
IRT calibration estimates item parameters from response data under a chosen model. Once items are calibrated, they can be used to estimate respondent locations on the same latent scale, subject to the linking and population assumptions of the measurement system.
The scale itself needs an origin and unit. These are conventional choices, often set by fixing a latent mean and variance in a reference population or using another identification convention.
A θ of 0 is not “zero ability”. It is a location on an arbitrarily scaled latent metric. Interpretation should be anchored to performance descriptors, item content and reference populations rather than treated as an absolute natural quantity.
15. Linking and Equating
Large assessment programs rarely administer exactly the same test forever. New items are introduced, old items retire and different forms may be used. Linking uses common items, common persons or other designs to place scores from different forms or cycles onto a common scale.
Equating is a stronger goal: making scores from different test forms interchangeable for a defined purpose. The validity of that exchange depends on design, content comparability, population assumptions, anchor stability and statistical method.
PISA’s technical documentation places major emphasis on scaling and maintaining trend comparability across cycles. The mathematical machinery matters because policy conclusions can depend on whether a change represents genuine proficiency movement or a break in the measurement link.
16. Differential Item Functioning
Differential item functioning, or DIF, occurs when respondents from different groups with the same modeled trait level have different probabilities of particular item responses.
DIF is a measurement warning, not automatic proof of unfairness. Sometimes the difference reveals unintended language, cultural or contextual demands. Sometimes it reflects a legitimate part of the construct. Domain review is needed to decide what the statistical flag means.
This is where psychometrics meets fairness. A scale can be mathematically precise yet substantively biased if items function differently for reasons outside the intended construct.
17. Computerised Adaptive Testing
Computerised adaptive testing uses an item bank and a provisional ability estimate to select the next item expected to be informative near the respondent’s current θ. After each response, the estimate is updated and another item is selected.
The attraction is efficiency: respondents need not answer many items that are far too easy or far too hard. The system can concentrate measurement where uncertainty is greatest.
But adaptive testing adds operational constraints. Content balance, item exposure, security, fairness, stopping rules and bank quality all matter. Maximising statistical information alone could produce a test that is psychometrically efficient but educationally unbalanced.
18. Person Scoring Is an Estimation Problem
Once item parameters are calibrated, a person’s θ can be estimated using maximum likelihood, Bayesian expected-a-posteriori methods, maximum-a-posteriori methods or related procedures.
Different scoring rules behave differently at the extremes and in short tests. A respondent who answers every item correctly or every item incorrectly can create difficulties for pure maximum-likelihood estimation in some simple IRT settings, which is one reason Bayesian or bounded approaches are commonly used.
The score is therefore an estimate with uncertainty, not a perfect coordinate.
19. Plausible Values and Population Reporting
Large-scale assessments often care more about population distributions and group comparisons than about producing a high-precision individual score for every respondent. In such settings, plausible values can represent uncertainty in latent proficiency while incorporating background information in a population model.
Plausible values are not simply several alternative test scores for an individual. They are draws designed for population-level inference under the assessment model. Treating them as interchangeable individual report-card scores misuses their purpose.
20. Model Fit Is Still Necessary
IRT can produce elegant curves even when the model is poorly matched to the responses. Fit should be examined at multiple levels: item fit, person fit where relevant, residual dependence, dimensionality, parameter stability and predictive behaviour.
OECD’s review of PISA scaling emphasises model-data fit and residual structures by sub-population. That reflects a core measurement principle: scaling is not complete when parameter estimation converges. The model must also be examined for where it systematically fails.
21. Sample Size and Parameter Stability
There is no single minimum sample size for IRT. Requirements depend on the model, number of items, item quality, trait distribution, response categories and precision demanded of the calibration.
Complex models such as 3PL or multidimensional IRT generally require more information than simpler Rasch or 1PL structures. Rare response categories in polytomous models can also make thresholds unstable. Simulation studies tailored to the intended design are often more informative than universal rules of thumb.
22. Missing, Omitted and Not-Reached Responses
A blank response can mean many things: skipped, not reached because time expired, intentionally omitted, inaccessible, technically missing or genuinely unknown. Treating every blank identically can distort item calibration and person scores.
Large assessment systems often distinguish response-status categories because speed, effort and administration conditions can create systematic missingness. A measurement model should not silently convert every absence into the same substantive state.
23. IRT Does Not Replace Content Validity
A mathematically excellent item bank can still measure the wrong thing. If a science assessment overweights reading complexity, the scale may partly measure language demand. If an advanced mathematics test omits a major part of the curriculum, highly precise θ estimates may still represent an incomplete construct.
Item statistics can tell us how questions behave. They cannot decide by themselves what the assessment ought to represent. That requires curriculum, domain and validity arguments.
24. IRT Does Not Make Scores Culture-Free
Because IRT models item behaviour, it can help detect group-specific functioning. But the model does not automatically remove cultural, linguistic or contextual influences. Those enter through item content, translation, opportunity to learn, test familiarity and the meaning of the construct itself.
International assessments therefore require both statistical scaling and substantive review. A common metric is a measurement achievement, not proof that every item means exactly the same thing everywhere.
25. Common Failure Modes
- Treating θ as an absolute natural quantity rather than a model-based scale location.
- Calling every lower-asymptote parameter “guessing” without examining item behaviour.
- Ignoring local dependence among passage-based or chained items.
- Assuming one latent dimension because the software was instructed to fit one.
- Using item parameters calibrated in one population as permanently invariant everywhere.
- Comparing groups without checking DIF or broader measurement invariance.
- Using adaptive item selection without protecting content balance and item exposure.
- Treating plausible values as individual scores.
- Equating forms that differ materially in construct or content.
- Believing a sophisticated scaling model can repair a weak assessment framework.
26. A Responsible IRT Workflow
- Define the construct and score use. Measurement begins with purpose.
- Design items for content coverage. Statistical efficiency cannot replace domain representation.
- Choose a model family deliberately. Rasch, 2PL, 3PL and polytomous models make different commitments.
- Check dimensionality and local dependence. Do not assume a one-factor item bank by default.
- Calibrate with sufficient data. Examine parameter uncertainty and stability.
- Inspect item fit and residuals. Convergence is not validation.
- Study DIF and subgroup behaviour. Fairness requires both statistical and substantive review.
- Protect scale links. Anchor items and equating designs need ongoing monitoring.
- Report measurement error. Precision varies across θ.
- Validate interpretation. Link the latent scale back to real tasks, descriptors and decisions.
27. How This Connects Across the eduKate Library
- How Construct Validity and Measurement Models Work — the broader owner for what a score means and whether the construct interpretation is defensible.
- How Measurement Invariance and Fair Comparisons Work — essential for group and longitudinal comparability.
- How Measurement Error and Misclassification Work — the wider route for noisy and imperfect measurement.
- How Statistical Power and Sample Size Planning Work — because calibration precision and hypothesis testing depend on design size and information.
- How Missing Data Analysis Works — especially relevant to omitted and not-reached assessment responses.
- How Structural Equation Modeling Works — a neighbouring latent-variable framework with measurement and structural components.
- How Principal Component Analysis and Factor Analysis Work — the route for exploratory and latent factor structure outside item-response modeling.
28. Authoritative Sources and Further Reading
- OECD, PISA 2022 Technical Report, published in its current report form by OECD, including the IRT scaling framework and trend methodology.
- OECD, Tomoya Okubo (2022), Theoretical Considerations on Scaling Methodology in PISA, OECD Education Working Papers No. 282.
- National Center for Education Statistics, NAEP Two-Parameter Logistic Model.
- National Center for Education Statistics, NAEP Three-Parameter Logistic Model.
- National Center for Education Statistics, NAEP Generalized Partial Credit Model.
- National Center for Education Statistics, NAEP Subject Area Scales and IRT scale documentation.
Final Idea
IRT changes the unit of thinking from “How many questions were correct?” to “What does each response tell us about a position on a defined measurement scale?” That extra sophistication can make assessments more precise, linkable and informative. It also raises the standard of responsibility: the latent scale, item behaviour, fairness, linking and score interpretation all have to remain answerable to the real construct the assessment claims to measure.