How Quasi-Experimental Designs Work | Natural Experiments, Difference-in-Differences, Regression Discontinuity, Instruments and Counterfactual Evidence

A school introduces a new programme in January. Scores rise by six points by June. Did the programme work?

Maybe. But the same six points could also be produced by maturation, easier assessment, a different cohort, a national policy, unusually strong teaching, recovery after a disruption or simple regression to the mean. The before-and-after difference is a change. The causal effect is the difference between what happened and what would have happened to the same units without the programme.

Quasi-experimental designs try to recover a credible counterfactual without assigning treatment by randomisation. They exploit rules, thresholds, timing, policy variation, natural events or carefully constructed comparison structures that make some untreated observations informative about what treated observations would otherwise have experienced.

This is not second-class experimentation by definition. Some quasi-experiments can support powerful causal claims. Others are observational comparisons with a confident label. The difference lies in the assignment mechanism and the assumptions that connect observed comparisons to the missing counterfactual.

The examples below are constructed. They do not describe eduKate programme effects or operational data.

The counterfactual is the missing object

For one unit at one time, we cannot observe both the outcome under treatment and the outcome under no treatment. The causal effect compares those potential outcomes, but one is unobserved. Research design determines which observed data can stand in for the missing state.

Randomised experiments use random assignment to make treatment status independent of potential outcomes in expectation. Quasi-experimental designs instead rely on a different source of comparability: a cutoff rule, common trend, valid instrument, sudden policy, untreated comparison series or another structured feature.

The World Bank’s Impact Evaluation in Practice treats instrumental variables, regression discontinuity and difference-in-differences as distinct routes for constructing counterfactuals. The design should be chosen from the assignment process, not from whichever estimator is fashionable.

Natural experiment and quasi-experiment are related, not identical

A natural experiment usually refers to a real-world event or rule that creates treatment variation resembling an experiment without the researcher assigning it. A quasi-experimental design is broader: it includes analytical designs that use non-random assignment but attempt to recover a valid causal comparison under explicit assumptions.

A lottery allocating scarce places can create near-random variation. A policy threshold can create a discontinuity. A staggered reform can create before-and-after comparisons across groups. Matching on observed characteristics is also often classified as quasi-experimental, although its causal credibility depends heavily on whether important unobserved differences remain.

Start with assignment, not the outcome graph

Before looking at results, ask how units came to receive treatment. Who decided? What rule? What timing? Could participants manipulate eligibility? Were some groups targeted because they were already improving or deteriorating? Did capacity limits create a queue or lottery?

The strongest quasi-experimental arguments often come from institutional detail. A statistical model cannot rescue an assignment process that systematically gives treatment to units already on a different trajectory unless the design has another valid source of identification.

Difference-in-differences: compare changes, not levels

Suppose one group receives a programme and another does not. Before the programme, their average outcomes are 60 and 70. Afterward, they are 75 and 78.

The treated group rose by 15. The comparison group rose by 8. Difference-in-differences attributes the extra 7-point rise to treatment under the relevant assumptions.

DiD estimate = (treated after − treated before) − (control after − control before).

The level gap of ten points at baseline is not automatically a problem. DiD can tolerate stable level differences. The critical question is whether the untreated trajectory of the treated group would have changed like the control group.

Parallel trends is a counterfactual assumption

Parallel trends does not mean the observed pre-treatment lines must be mathematically identical. It means that, absent treatment, the relevant average changes would have been comparable in the way required by the estimand and model.

Pre-treatment data can challenge implausible designs. If the treated group was already accelerating upward while controls were flat, a simple parallel-trends story looks weak. But similar observed pre-trends do not prove the unobserved post-treatment counterfactual would remain parallel.

The World Bank’s discussion of difference-in-differences highlights counterfactual parallel trends as the key assumption. Modern work also shows that staggered treatment timing and heterogeneous treatment effects can make simple two-way fixed-effects estimators difficult to interpret.

Staggered adoption changes the DiD problem

When different groups adopt a policy at different times, already-treated units can become implicit controls for later-treated units in conventional regressions. If treatment effects vary over time, such comparisons can receive unexpected weights and produce an estimate that does not equal a transparent average treatment effect.

Recent econometric work has developed estimators that define group-time effects more explicitly and use cleaner comparison sets. NBER surveys published through 2025–2026 discuss heterogeneous treatment effects, event-study normalisation, pre-trend testing and modern alternatives to naïve two-way fixed-effects practice.

The lesson for non-specialists is simple: “we used fixed effects” is not enough. Ask who is being compared with whom at each time.

Event studies show dynamics but can create false comfort

An event study aligns observations around treatment timing and estimates effects before and after the event. Pre-event coefficients can reveal obvious differential trends; post-event coefficients can show whether effects emerge, grow, fade or reverse.

But “no significant pre-trend” is not proof of parallel trends. Pre-treatment estimates can be imprecise. Testing many pre-periods creates multiplicity. Conditioning interpretation on passing a pre-trend test can itself distort inference.

Use pre-periods to understand the design, not as a ritual certificate.

Regression discontinuity: compare units around a cutoff

Imagine a support programme offered to learners scoring below 50 on a placement test. A learner scoring 49 is eligible; one scoring 51 is not. If units cannot precisely manipulate the score and other determinants of outcome vary smoothly through the cutoff, outcomes immediately around 50 can provide a local causal comparison.

The treatment assignment changes abruptly at the threshold while many background characteristics should not. A jump in outcome at the same threshold can therefore be attributed to treatment under the design assumptions.

The World Bank’s Regression Discontinuity resource describes the continuous assignment variable, cutoff and local nature of the effect. The What Works Clearinghouse also maintains current advanced group-design standards that include regression discontinuity designs.

The RDD effect is usually local

If the design is credible near score 50, it does not automatically estimate the effect for a learner scoring 20 or 90. The counterfactual comparison is strongest around the threshold where treated and untreated units are similar except for eligibility.

This local nature is not a weakness if the decision concerns the threshold population. It becomes a problem only when a local effect is marketed as universal.

Manipulation can destroy a threshold design

If staff can nudge scores below 50 to obtain programme access, units on either side of the cutoff may no longer be comparable. Sorting around the threshold becomes part of the treatment assignment process.

Researchers therefore inspect the density of the running variable, institutional rules and covariate continuity. No single diagnostic proves the design valid. The institutional story and statistical evidence should agree.

Bandwidth is a bias–variance decision

Using observations very close to the cutoff improves local comparability but reduces sample size. Expanding the bandwidth provides more data but asks the functional form to model a wider region where treated and untreated units may differ more.

A responsible RDD reports how results respond to reasonable bandwidth and specification choices without treating a preferred bandwidth as the only possible universe.

Sharp and fuzzy discontinuities answer different compliance problems

In a sharp design, crossing the threshold deterministically changes treatment. In a fuzzy design, it changes the probability of treatment. Some eligible units decline; some ineligible units obtain access through another route.

The threshold can then act as an instrument for actual treatment. Interpretation shifts toward a local effect for units whose treatment status is changed by threshold eligibility under additional assumptions.

Instrumental variables: use variation that moves treatment through a valid channel

An instrument Z must influence treatment X and, for the intended causal interpretation, affect outcome Y only through the treatment pathway relevant to the estimand while satisfying the required independence assumptions.

Examples can include random encouragement, distance created by administrative rules or eligibility thresholds. Merely finding a variable strongly correlated with treatment does not make it a valid instrument.

The difficult condition is often exclusion: could the instrument affect the outcome through another channel? A scholarship eligibility rule might change motivation or expectations even among people who do not take up the scholarship, complicating the pathway.

Weak instruments can produce unstable answers

If the proposed instrument barely changes treatment, the causal estimate can become extremely noisy and sensitive to specification. A large sample does not necessarily solve a fundamentally weak first stage.

Instrument strength should be examined alongside the institutional argument. A strong association without credible exclusion is not enough; credible exclusion with almost no treatment movement may also be uninformative.

LATE: the instrument may identify an effect for compliers, not everyone

Under standard monotonicity and exclusion assumptions, a binary instrument can identify a local average treatment effect for units whose treatment is changed by the instrument—the compliers.

If encouragement changes participation only among moderately interested people, the estimated effect need not equal the effect among people who would always participate or never participate. External validity becomes part of interpretation.

This connects directly to External Validity and Evidence Transfer.

Interrupted time series: ask whether the level or trend changes at an intervention point

An interrupted time series uses many observations before and after a clearly timed intervention. Instead of comparing one before point with one after point, it estimates the pre-intervention level and trend, then asks whether the post-intervention series departs from that projected trajectory.

This can be powerful when a policy affects an entire population at once. It is vulnerable when another event occurs at the same time, measurement changes, seasonality is ignored or the number of time points is too small to estimate the underlying trend credibly.

Controlled interrupted time series strengthens the counterfactual

A comparison series not exposed to the intervention can help separate the policy from common shocks. If both regions change at the intervention date, the event may not be the cause. If only the treated region shows the break, the design becomes more informative, assuming the control series is genuinely comparable.

This begins to resemble difference-in-differences with richer time structure.

Synthetic control builds a comparison from weighted untreated units

When one city, country or organisation receives a policy, no single untreated unit may look comparable. Synthetic control methods choose weights over several untreated units so their pre-treatment outcomes and predictors resemble the treated unit.

The post-treatment divergence between the treated unit and its synthetic comparison can then be interpreted causally under the design assumptions. Placebo analyses across untreated units can show whether the treated gap is unusual relative to gaps that appear when treatment is falsely assigned elsewhere.

The design is most persuasive when pre-treatment fit is strong, treatment is not anticipated in ways that alter earlier outcomes, and there are enough suitable donor units.

Matching balances observed characteristics, not hidden ones

Matching pairs or weights treated and untreated units with similar observed characteristics. Propensity-score approaches summarise observed treatment probabilities under a model. These techniques can improve comparability on what was measured.

They do not create randomisation. If an unmeasured motivation variable influences both programme uptake and outcome, matching on age, baseline score and school may leave serious confounding.

Matching is therefore strongest when treatment selection is plausibly captured by rich pre-treatment information and the overlap between groups is good. It is weaker when treatment decisions depend on unrecorded judgement or self-selection.

Propensity scores are not quality scores

A propensity score estimates the probability of treatment conditional on measured covariates. A score of 0.8 does not mean an observation is “80% matched” or that causal bias has been removed.

After matching or weighting, researchers should examine balance in the underlying covariates and inspect regions without common support. A single model diagnostic should not replace design reasoning.

Negative controls can probe some hidden biases

A negative-control outcome is one that the treatment should not plausibly affect but that may share confounding structure with the target outcome. A negative-control exposure works in the opposite direction.

Unexpected associations can reveal problems. A null result does not guarantee absence of confounding. The control is useful only if the assumptions making it negative are credible.

Placebo interventions test the design against events that did not happen

In a time-series or DiD design, researchers may assign a false treatment date before the real intervention. A large “effect” at the placebo date suggests the model is detecting background trends or specification artefacts rather than the intervention.

Placebos are diagnostics. Passing them does not prove the real design valid. Failing them provides useful evidence that the design needs repair.

Spillovers can contaminate the comparison group

If treated learners share materials with untreated classmates, controls are no longer fully untreated. A road improvement in one district can redirect traffic from neighbouring districts. A public information campaign can cross administrative borders.

Quasi-experimental designs often rely on a stable untreated comparison. Spillovers violate that simple story. Researchers may need to redefine treatment exposure, model interference or select comparison areas less likely to be affected.

Anticipation changes the pre-treatment period

If schools know a policy will start next term, they may alter staffing or behaviour before the official date. Those periods are no longer clean untreated observations.

Event-study plots may show a “pre-trend” that is actually early response to anticipated treatment. The assignment and information timeline must therefore be reconstructed from real institutional events, not just database dates.

Concurrent reforms are the enemy of simple attribution

Suppose a new curriculum and a teacher training programme begin together. A time-series break after implementation cannot tell which component caused the change unless another design feature separates them.

Calling the package “the treatment” can be honest if the causal question concerns the package. It becomes misleading when the result is attributed to one component that was never isolated.

Measurement changes can manufacture discontinuities

If an outcome definition changes exactly when a policy begins, an apparent effect may be a measurement break. Administrative systems often revise coding rules, eligibility definitions or reporting standards over time.

Check the measurement timeline before interpreting the outcome timeline. See Measurement Error and Misclassification for the dedicated owner.

Quasi-experiments often estimate local or design-specific effects

RDD estimates near a cutoff. Instrumental variables may estimate effects for compliers. A synthetic control may answer a question about one treated region. DiD estimates depend on treated groups, timing and comparison sets.

The effect’s domain should therefore travel with the headline. “The programme increased outcomes by five points” is incomplete if the design identifies a five-point effect only among threshold-near applicants in one policy period.

Inference must respect the number of independent policy changes

A dataset can contain millions of individual rows but only three independently treated regions. Standard errors based on individual-level sample size can then be badly misleading.

Policy designs often require clustering at the level of assignment or methods suited to a small number of treated clusters. The effective information is constrained by the treatment variation, not merely the number of observations stored.

This connects to Clustered and Multilevel Data.

Specification searching can create a quasi-experiment after the fact

Researchers can try many bandwidths, control groups, time windows, covariates and functional forms. If only the favourable specification is reported, design uncertainty disappears from view.

Pre-specification where feasible, transparent robustness analysis and reporting of reasonable alternatives help protect interpretation. The goal is not to freeze research into one rigid model; it is to stop researcher flexibility from masquerading as inevitability.

See Sensitivity Analysis and Robustness Checks.

A design can be strong even when the effect is imprecise

A credible discontinuity with few units near the cutoff may produce a wide confidence interval. That is not the same as a weak causal design. It is a strong identification argument with limited precision.

Conversely, a huge confounded dataset can produce an extremely precise estimate of the wrong quantity. Precision and causal credibility are separate axes.

The design diagram

CAUSAL QUESTION
→ DEFINE TREATMENT + COMPARATOR + OUTCOME + TIME
→ RECONSTRUCT ASSIGNMENT PROCESS
→ IDENTIFY NATURAL SOURCE OF COMPARABILITY
→ CHOOSE DESIGN
→ STATE IDENTIFICATION ASSUMPTIONS
→ TEST OBSERVABLE IMPLICATIONS
→ ESTIMATE TARGET EFFECT
→ RUN DESIGN-SPECIFIC ROBUSTNESS CHECKS
→ DEFINE LOCAL / POPULATION DOMAIN
→ REPORT UNCERTAINTY
→ OBSERVE WORLD RETURN

The estimator is only one step in the middle. The design starts before the regression and continues after the coefficient is printed.

A practical design audit

When quasi-experimental evidence deserves more trust

When the label should not persuade you

Method names are not causal credentials.

The relationship to experiments

Randomised and quasi-experimental designs share the same underlying problem: the missing counterfactual. Randomisation solves part of that problem by construction. Quasi-experiments solve it through structured assumptions about naturally occurring variation.

The dedicated experiment owner remains How Experimental Design Works. This article owns the non-randomised identification designs that require a different inferential bridge.

The relationship to observational studies

All quasi-experiments are observational in the sense that treatment is not randomly assigned by the researcher. Not all observational studies are quasi-experiments. The quasi-experimental label is most useful when the design exploits a specific assignment mechanism or natural comparison that goes beyond generic statistical adjustment.

See How Observational Studies Work for cohort, cross-sectional and general non-randomised evidence.

The relationship to longitudinal data

Difference-in-differences, interrupted time series and event studies often require repeated observations. The newly established How Longitudinal and Panel Data Work article owns repeated-state structure, attrition, within-unit variation and temporal measurement. This article owns the causal designs built on particular patterns inside such data.

A worked interpretation

Suppose a fictional DiD estimate is +7 points. A responsible summary is not “the programme caused a seven-point gain” without qualification. It is: “Under the assumption that the treated group’s untreated outcome would have followed the comparison group’s change over the same period, the estimated average treatment effect for the treated units in this design is seven points.”

That sentence is longer. It is also reusable. Another researcher can inspect the comparison group, pre-trends, timing and treatment domain rather than treating the number as a detachable fact.

The World Return of quasi-experimental evidence

Many important questions cannot be randomised. Governments cannot randomly assign recessions. Cities cannot rerun infrastructure histories. Schools may have policies already implemented before an evaluation begins.

Quasi-experimental methods allow researchers to learn from these imperfect worlds without pretending the imperfections disappeared. Their contribution is not statistical magic. It is disciplined comparison: finding where the world itself created a useful contrast, then being precise about what that contrast can and cannot support.

Sources and current methodological routes

Sources were checked for this edition on 5 September 2026. The links below support method definitions and current developments; the numerical education examples are original hypothetical illustrations.

Continue through eduKate

Wintour return: The question is never whether a dataset “looks experimental”. The question is whether the mechanism that assigned treatment created a comparison strong enough to carry the causal claim being asked of it.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading