“X happened before Y” is not enough. “X and Y move together” is not enough. “People exposed to X have more Y” is not enough. Causal inference begins when we ask a harder question: what would have happened to the same relevant system if the exposure or action had been different?
That alternative world is usually unobservable. A student either received an intervention or did not. A city either changed a transport policy at that time or it did not. A company either adopted a new process or continued with the old one. We observe one realised path, not both.
Causal inference is the discipline of using design, assumptions and evidence to compare the observed world with a credible counterfactual. Randomised experiments are powerful because assignment can make treatment groups comparable by design. Observational research can also support causal inference when the structure creating exposure variation is understood well enough and the assumptions are explicit.
The World Bank’s Impact Evaluation in Practice builds causal evaluation around this same challenge: identify a credible counterfactual for what would have happened without the intervention. Harvard T.H. Chan School of Public Health’s CAUSALab develops and teaches modern methods for causal inference in randomised and observational settings. The shared principle is simple: a causal estimate is not produced by a statistical command. It is produced by a research design whose assumptions make the comparison meaningful.
The causal inference loop
CAUSAL QUESTION → INTERVENTION / EXPOSURE → OUTCOME → TARGET POPULATION → COUNTERFACTUAL → CAUSAL GRAPH / ASSUMPTIONS → DESIGN → DATA → IDENTIFICATION → ESTIMATION → DIAGNOSTICS → SENSITIVITY → TREATMENT EFFECT → DECISION → NEW EVIDENCE / REPLICATION
Every step matters. If the counterfactual is poorly defined, a precise estimate can answer the wrong question. If the design does not identify the effect, more sophisticated estimation cannot rescue it.
1. Causality is about interventions, not just associations
An association asks whether two variables differ together. A causal question asks what would change if one variable were intervened upon while relevant alternatives were held or accounted for.
Students who sleep more may score better. That association could arise because sleep improves learning, because organised students both sleep and study better, because illness reduces both sleep and scores, or because other factors affect both. The causal question requires us to separate these pathways.
2. The counterfactual is the missing comparison
For one unit, we can imagine an outcome under treatment and an outcome under no treatment. Only one is observed. The other is counterfactual.
Causal inference therefore relies on groups, time, rules, experiments or models to approximate the missing comparison. The quality of the estimate depends on whether that approximation is credible for the target question.
3. Potential outcomes formalise the causal question
In the potential-outcomes framework, each unit has conceptual outcomes under different treatment states. The causal effect for that unit is the difference between those potential outcomes.
Because both potential outcomes cannot normally be observed for the same unit at the same time, causal effects are estimated across populations under assumptions that make observed groups valid stand-ins for each other’s missing outcomes.
4. The estimand must be specified before the estimator
An estimand is the causal quantity we want to know. Common examples include:
- Average Treatment Effect (ATE): average effect in the target population.
- Average Treatment Effect on the Treated (ATT): average effect among units that actually received treatment.
- Average Treatment Effect on the Untreated (ATU): effect among units that did not receive treatment.
- Local Average Treatment Effect (LATE): effect for a specific subgroup identified by an instrumental-variable design.
- Conditional effect: effect for units with particular characteristics.
The estimator—regression, matching, weighting, difference-in-differences or another method—should be chosen after the estimand and identification strategy are clear.
5. Treatment must be well defined
“Tutoring”, “policy”, “exercise”, “technology use” and “management quality” can contain many versions. If different units receive materially different forms of the treatment, the causal question becomes ambiguous.
A well-defined intervention specifies dose, timing, duration, delivery and relevant alternatives. “One 90-minute small-group lesson each week for twelve weeks” is more causally interpretable than “received tutoring”.
6. Outcomes need timing and measurement rules
The same intervention can have different short-term and long-term effects. A programme may improve immediate test performance but not retention, or create delayed benefits that are invisible after one week.
Causal questions should therefore specify the outcome, measurement method and time after intervention.
7. The target population determines what the effect means
An effect estimated among highly selected volunteers may not equal the effect in the broader population. A policy effect in one city may not transport to another with different infrastructure or institutions.
Every causal estimate should therefore be attached to a population, place and period rather than presented as a universal constant.
8. Randomisation creates comparability in expectation
Random assignment breaks systematic relationships between treatment and pre-treatment characteristics in expectation. This is why randomised experiments can estimate causal effects with relatively few structural assumptions.
Randomisation does not guarantee perfect balance in every finite sample, eliminate missing data, prevent non-compliance or make outcome measurement unbiased. It protects the assignment mechanism; the rest of the study still needs quality control.
See How Experimental Design Works for randomisation, blocking, blinding and experimental execution.
9. Randomisation answers the effect of assignment first
In trials with imperfect compliance, the cleanest causal contrast is often the effect of being assigned to treatment—the intention-to-treat effect. The effect of actually receiving treatment can require additional assumptions because compliance itself may be related to outcomes.
This distinction prevents researchers from quietly replacing a randomised comparison with a self-selected one after the study begins.
10. Observational causal inference needs exchangeability
In observational data, treatment groups often differ before treatment. Causal inference requires a form of conditional comparability: after accounting for the right pre-treatment variables, the treated and untreated groups must be sufficiently exchangeable for the causal contrast.
This is the difficult assumption commonly described as no unmeasured confounding under the relevant conditioning set.
See How Observational Studies Work for the study-design layer beneath naturally occurring variation.
11. Confounding mixes treatment selection with outcome risk
A confounder influences both treatment assignment and the outcome. If not handled appropriately, it can make treatment groups look different because of who entered them rather than because of treatment itself.
Students with greater difficulties may receive more support. Sicker patients may receive more intensive treatment. High-risk neighbourhoods may receive more safety interventions. In each case, naive outcome comparison can make helpful interventions appear harmful.
12. Confounding is a causal concept, not a correlation threshold
A variable is not a confounder merely because it correlates with treatment and outcome in the observed dataset. Its role depends on the causal structure.
This is why automated “adjust for every significant variable” procedures are dangerous. Variables can be confounders, mediators, colliders, instruments, proxies or consequences of treatment. Conditioning on the wrong type can introduce bias.
13. Directed acyclic graphs make assumptions visible
A directed acyclic graph, or DAG, represents hypothesised causal arrows among variables. It is not proof of the causal structure. It is a way to expose assumptions so they can be inspected.
DAGs help answer questions such as:
- Which variables confound the treatment-outcome relationship?
- Which variables lie on the causal pathway?
- Which variables are colliders that should not be conditioned on?
- Which pathways should be blocked to identify the desired effect?
- Which variables can remain unmeasured without biasing the target contrast?
14. Mediators are part of the mechanism
If treatment affects a mediator which then affects the outcome, controlling for that mediator changes the estimand. It can remove part of the total effect.
For example, if a teaching intervention improves study habits and those habits improve examination performance, adjusting for post-intervention study habits can remove a pathway through which the intervention works.
15. Colliders can create associations by conditioning
A collider is influenced by two variables. Conditioning on it can open a non-causal pathway and create an association between otherwise unrelated causes.
This is one of the most counterintuitive lessons in causal inference: adding more control variables can make an analysis worse.
16. Positivity requires real comparison
For every combination of covariates relevant to the target population, there must be some possibility of receiving each treatment condition being compared. If certain students always receive one intervention and never another, data cannot reveal the alternative outcome for those students without extrapolation.
Positivity violations appear as lack of overlap. Extreme propensity scores are a warning that the comparison is leaning on model assumptions rather than observed support.
17. Consistency links observed treatment to the intervention definition
Consistency requires that the observed outcome under the treatment actually received corresponds to the potential outcome under that well-defined treatment condition.
If “treatment” contains many materially different versions, consistency becomes difficult. This returns us to the need for precise intervention definitions.
18. Interference means one unit’s treatment affects another unit
Many causal models assume one unit’s outcome depends only on its own treatment. That fails in infectious disease, classrooms, social networks, markets, transport and policy systems where treatment spills over.
A tutoring programme can change peer interactions. A vaccination programme can protect untreated people. A road closure reroutes traffic onto neighbouring streets. These are not nuisance effects; they are part of the system.
19. Spillovers change the estimand
When interference exists, researchers may need to estimate direct effects, indirect effects, total effects or effects under different treatment coverage levels.
The causal question becomes networked: “What happens to this unit when it is treated and when different fractions of neighbouring units are treated?”
20. Regression adjustment estimates conditional contrasts
Regression can adjust for observed pre-treatment differences and estimate treatment effects under specified models. Its causal validity depends on the conditioning set and model assumptions, not on the existence of a coefficient.
Flexible models can reduce functional-form error, but no regression can adjust for an unmeasured confounder whose effect is absent from the data.
21. Standardisation predicts counterfactual outcomes
G-computation or standardisation models the outcome under different treatment values and averages those predictions over the target population. Conceptually, it asks each observed unit: what would your outcome be predicted to look like under treatment and under no treatment?
The method relies on correct modelling of the relevant conditional outcome process and the same identification assumptions as other observational approaches.
22. Propensity scores model treatment assignment
The propensity score is the probability of receiving treatment given observed pre-treatment covariates. Units with similar propensity scores have similar measured treatment-prediction profiles.
Propensity methods can support matching, weighting, stratification or covariate adjustment. Their purpose is balance on observed confounders, not prediction accuracy for its own sake.
23. Balance matters more than propensity-model fit
A propensity model can predict treatment well while producing poor overlap or balance. The diagnostic question is whether treated and comparison groups become comparable on measured pre-treatment variables after the adjustment strategy.
This is a good example of method purpose: the score is a tool for creating comparability, not an end in itself.
24. Inverse probability weighting creates a synthetic target population
Inverse probability weighting gives more weight to units whose treatment was less expected given their covariates. Under the assumptions, this creates a weighted population in which treatment is less associated with measured confounders.
Extreme weights reveal poor overlap and can make estimates unstable. Trimming or redefining the target population may be more honest than forcing unsupported comparisons.
25. Doubly robust methods combine two models
Doubly robust estimators combine an outcome model with a treatment-assignment model. Under suitable conditions, the causal effect can remain consistently estimated if one of the two nuisance models is correctly specified.
This robustness is valuable but not magical. Both models still rely on measured covariates and the underlying identification assumptions.
26. Matching creates local comparability
Matching pairs or groups units that look similar on selected pre-treatment characteristics. It can make the comparison more transparent and reduce model extrapolation.
The trade-off is that unmatched units may be discarded, changing the target population. Matching quality should be assessed through balance and overlap, not merely match count.
27. Exact matching reveals the curse of dimensionality
Matching exactly on many variables quickly becomes impossible because few units share identical profiles. Propensity scores, distance metrics and coarsening reduce dimensionality, but they also introduce modelling choices.
Causal design is therefore often a compromise between comparability and available support.
28. Difference-in-differences uses changes as the counterfactual
Difference-in-differences compares how outcomes change over time in a treated group relative to a comparison group.
The key identifying assumption is a form of parallel trends: absent treatment, the groups would have followed sufficiently similar outcome trends. This assumption cannot be proven from post-treatment data and should be investigated using pre-treatment patterns and institutional knowledge.
29. Parallel pre-trends support but do not prove parallel counterfactual trends
Similar historical trends make the design more plausible, but a new shock coinciding with treatment can still break comparability. Researchers should search for co-occurring policies, composition changes and anticipatory behaviour.
30. Event studies reveal dynamics around an intervention
Event-study designs estimate effects at multiple times before and after an intervention. They can show anticipation, delayed effects and persistence.
When treatment timing differs across units, naive two-way fixed-effects implementations can behave poorly under heterogeneous treatment effects. Modern designs require careful cohort and timing treatment.
31. Regression discontinuity uses an assignment threshold
When treatment changes sharply at a cutoff, units just above and below the threshold may be similar enough to support a local causal comparison.
Researchers should test whether other variables change discontinuously at the threshold, whether units can manipulate their assignment variable and whether the effect is local to the cutoff rather than universal.
32. Instrumental variables use external leverage
An instrument affects treatment but, under the identification assumptions, affects the outcome only through treatment and is otherwise sufficiently independent of relevant outcome causes.
The exclusion restriction is a causal assumption about the world. It cannot be established by a strong first-stage statistical relationship alone.
33. Weak instruments create unstable causal estimates
If an instrument barely changes treatment, the causal estimate can become noisy and sensitive. Strong treatment prediction is necessary but not sufficient; validity remains the harder requirement.
34. Instrumental-variable effects can be local
Under common assumptions, instrumental variables identify an effect for “compliers”—units whose treatment status changes because of the instrument. That local average treatment effect may differ from the effect in the entire population.
Estimand interpretation therefore matters as much as estimation.
35. Synthetic control constructs a missing comparison trajectory
When one unit receives a major intervention, a weighted combination of untreated units can sometimes reproduce its pre-intervention trajectory. That synthetic control becomes a candidate counterfactual for the post-intervention period.
The design is strongest when pre-treatment fit is good, the donor pool is credible and no simultaneous shock uniquely affects the treated unit.
36. Interrupted time series asks whether the process changed
An interrupted time-series design examines changes in level or trend after an intervention. It can be useful when many repeated observations exist before and after a policy change.
The major threat is another event occurring at the same time. A long pre-intervention series helps model existing trend and seasonality but does not eliminate concurrent causes.
37. Target trial emulation makes observational questions more explicit
Harvard causal-inference work has popularised the idea of specifying the hypothetical randomised trial that an observational study is trying to emulate. That means defining eligibility, treatment strategies, assignment procedure, follow-up, outcome, causal contrast and analysis plan.
The exercise reveals hidden ambiguities. If the hypothetical trial cannot be described clearly, the observational causal question may not be sufficiently defined.
38. Time-varying treatment complicates ordinary adjustment
In longitudinal systems, treatment can change over time and be affected by prior outcomes or confounders. Those confounders may themselves be affected by earlier treatment.
Ordinary adjustment can then block part of the effect or create bias. G-methods such as marginal structural models were developed for these treatment-confounder feedback structures.
39. Mediation asks how an effect travels
Mediation analysis separates pathways through intermediate variables. It asks whether treatment changes the outcome partly through a mediator and how much effect may remain through other pathways.
Mediation requires stronger assumptions than total-effect estimation because mediator-outcome confounding can be difficult to control, especially when affected by treatment.
40. Mechanism evidence and effect estimation complement each other
Average treatment effects tell us how outcomes changed under a treatment contrast. Process tracing, qualitative research and mechanistic evidence can help explain why.
See How Case Study Research and Process Tracing Work. Causal inference is strongest when population-level effect estimates and within-case mechanism evidence constrain one another.
41. Heterogeneous effects mean the average can hide different realities
An intervention can help one subgroup, have little effect on another and harm a third while the average looks modestly positive.
Subgroup analysis should be theory-driven and adequately supported by data. Searching many subgroups after seeing results can manufacture apparent heterogeneity by chance.
42. Individual treatment effects are harder than group averages
Estimating the average causal effect does not mean we know the effect for a particular person or unit. Individual counterfactual outcomes remain unobserved.
Personalised-treatment models require much stronger information and validation than population averages. High predictive accuracy for outcomes does not automatically imply accurate individual causal effects.
43. Machine learning can estimate nuisance functions flexibly
Modern causal methods use machine learning to estimate treatment probabilities, outcome functions and conditional effects. Cross-fitting and orthogonalisation can reduce overfitting bias in high-dimensional settings.
Machine learning improves estimation flexibility. It does not create identification. If an essential confounder is missing or the causal graph is wrong, a more powerful predictor can estimate the wrong causal quantity with greater precision.
44. Causal discovery is not causal identification
Algorithms can search for graph structures consistent with statistical dependencies. These methods can generate hypotheses and narrow possibilities under assumptions.
They should not be confused with proving the direction of real-world causation from observational correlations alone. Domain knowledge, design and interventions remain essential.
45. Negative controls test for hidden structure
A negative-control outcome should not plausibly be caused by the treatment. A negative-control exposure should not plausibly cause the target outcome. Unexpected associations can indicate residual confounding, selection or measurement bias.
Negative controls turn some untestable concerns into observable diagnostics.
46. Placebo tests ask whether effects appear where they should not
Policy and quasi-experimental research often uses placebo dates, placebo groups or placebo outcomes. If the method “finds” an effect before the intervention or in unaffected groups, the identification strategy may be capturing unrelated structure.
47. Sensitivity analysis measures fragility to hidden confounding
No observational analysis can simply assert that unmeasured confounding does not exist. Sensitivity analysis asks how strong hidden confounding would need to be to explain away or materially change the estimated effect.
This shifts the conversation from impossible certainty to quantified robustness.
48. Specification curves can reveal analytic flexibility
When many reasonable analytical choices exist, researchers can examine how estimates change across specifications rather than highlighting only one preferred model.
The exercise is useful only when the specification set represents genuinely defensible alternatives rather than arbitrary model fishing.
49. Causal effects can change through time
An intervention introduced during early adoption may have a different effect after institutions learn, infrastructure matures or behaviour adapts. Historical causal estimates are not permanent constants.
Freshness and context matter. A treatment effect should be versioned conceptually by population, implementation and time.
50. Transportability asks whether effects move to new populations
A valid effect in one study population may not apply directly elsewhere. Baseline risk, implementation quality, institutions, culture, geography and treatment versions can differ.
Transportability requires identifying which effect modifiers differ between source and target populations and whether enough overlap exists to support reweighting or reasoned transfer.
See How Comparative Systems Research Works.
51. Policy effects include implementation
A policy label can hide substantial variation in enforcement, funding, take-up and local adaptation. Estimating the effect of “the policy” without measuring implementation can mix treatment concept with delivery quality.
Causal evaluation should distinguish assignment, uptake, fidelity and exposure intensity where those differences matter.
52. Causal evidence should not be reduced to one hierarchy
Randomised trials provide strong causal identification for many questions, but not every important intervention can be randomised, and a poorly executed trial can still be weak. Natural experiments, longitudinal studies, mechanism evidence and replicated observational designs can all contribute.
The right question is not “Which design has the highest label?” but “Which design most credibly identifies this causal contrast under these constraints?”
53. Systematic reviews need causal compatibility
Combining studies requires more than matching outcome names. Studies may estimate different treatment versions, populations, follow-up periods or causal estimands.
See How Systematic Reviews and Evidence Synthesis Work. Meta-analysis can increase precision only when the combined effects are meaningfully commensurable.
54. Statistical significance does not establish causality
A p-value concerns the observed data under a statistical model. It does not tell us whether the treatment groups were exchangeable, whether an instrument was valid, whether a collider was conditioned on or whether the causal estimand was well defined.
Statistical inference and causal identification are different layers. See How Statistical Inference and Uncertainty Work.
55. Precision cannot repair identification failure
A confidence interval can be extremely narrow around a badly biased estimate. Big data reduces random error while leaving systematic error intact.
The first question is whether the effect is identified. Only then does precision become meaningful.
56. A causal graph and an estimation model do different jobs
The causal graph encodes assumptions about relationships in the world. The estimation model maps observed data into an effect estimate. A highly flexible estimator cannot compensate for a mistaken causal graph.
Keeping these layers separate makes analysis easier to challenge and improve.
57. Data provenance belongs inside causal inference
Variables may be derived from administrative rules, sensors, surveys or historical records. A change in coding, eligibility, measurement or data cleaning can create apparent treatment or outcome changes.
Causal analysts therefore need the same provenance discipline as data engineers. See How Data Management Works and Data Quality.
58. Causal inference needs a pre-analysis causal story
Researchers should articulate treatment, outcome, timing, confounders, mediators and plausible alternative pathways before inspecting every possible association. This reduces the temptation to construct a causal story after seeing the result.
59. Pre-registration cannot make bad assumptions good
Pre-registration preserves which hypotheses and analyses were specified before outcomes were inspected. It reduces undisclosed flexibility but does not validate the causal model, treatment definition or data quality.
Transparency is necessary for trust; it is not a substitute for good design.
60. Causal claims should carry an assumption receipt
A mature causal publication should make visible:
- target population;
- treatment definition;
- outcome definition and timing;
- causal estimand;
- assignment or identification mechanism;
- adjustment set;
- overlap or positivity diagnostics;
- missing-data assumptions;
- interference assumptions;
- estimation method;
- sensitivity analyses;
- effect estimate and uncertainty;
- known limits to transportability.
61. A causal inference checklist for learners
- What intervention or exposure is being compared?
- What is the outcome?
- What population does the effect refer to?
- What is the counterfactual?
- How were treatment groups created?
- What confounders affect both treatment and outcome?
- Are any adjusted variables mediators or colliders?
- Is there enough overlap between groups?
- Can one unit’s treatment affect another?
- What assumptions identify the effect?
- What method estimates the effect?
- What sensitivity checks were performed?
- How large is the effect and how uncertain is it?
- Does the effect travel to the population I care about?
62. A causal design workflow
QUESTION → DEFINE TARGET TRIAL / INTERVENTION → TARGET POPULATION → TREATMENT STRATEGIES → OUTCOME + TIME → ESTIMAND → CAUSAL GRAPH → IDENTIFICATION ASSUMPTIONS → DESIGN → DATA PROVENANCE → OVERLAP CHECK → ESTIMATION → DIAGNOSTICS → NEGATIVE / PLACEBO TESTS → SENSITIVITY → HETEROGENEITY → TRANSPORTABILITY → DECISION → REPLICATION
63. Where causal inference sits in the eduKate Library
Causal inference is a methodological bridge among several existing owners. Experimental Design owns controlled assignment. Observational Studies owns naturally occurring variation and its bias structure. Statistical Inference owns estimation and uncertainty. Case Study Research and Process Tracing owns within-case mechanism evidence.
This article owns the cross-domain question connecting them: under what assumptions can an observed contrast be interpreted as the effect of changing something?
Domain-specific causal claims in Biology, Medicine and Veterinary remain owned by those domains. They can route here for generic causal machinery without merging their scientific ownership.
64. World Return from causal knowledge
Description tells us what is happening. Prediction tells us what may happen next. Causal inference gives us something different: evidence about what may happen if we change the system.
That makes causal knowledge unusually consequential. It can support policy, education, engineering, operations and science—but only if the assumptions survive scrutiny. A bad causal claim can make an institution confidently intervene in the wrong mechanism.
65. The deepest causal habit is to protect the alternative world
Every causal claim depends on a world that did not happen. We cannot observe it directly, so the research design must construct a credible substitute.
Randomisation, matching, natural experiments, longitudinal comparisons, instruments and causal models are different ways of protecting that missing comparison from contamination. The strongest causal work never forgets that the counterfactual is an inference, not an observation.
Sources and further reading
- World Bank — Impact Evaluation in Practice
- World Bank — ImpactAI causal-evidence platform
- Harvard T.H. Chan School of Public Health — CAUSALab
- STROBE — Reporting Observational Studies
- IARC / NCBI Bookshelf — Confounding and Bias Assessment
Continue through eduKate
- How Research Methods and Source Evaluation Work
- How Experimental Design Works
- How Observational Studies Work
- How Statistical Inference and Uncertainty Work
- How Case Study Research and Process Tracing Work
- How Systematic Reviews and Evidence Synthesis Work
- How Comparative Systems Research Works
Final idea: Causal inference is the art of making one missing world comparable to the world we observed. Its credibility comes from the design and assumptions that protect that comparison—not from the sophistication of the final equation.