How External Validity and Evidence Transfer Work | When Research Travels to a New Population

A carefully conducted study finds that a programme improves an outcome by almost seven points. Another institution adopts the programme and obtains an improvement closer to three. The immediate temptation is to decide that one institution implemented it badly, that the original research was exaggerated, or that research cannot be trusted.

There is another possibility: both results can be compatible with the same underlying pattern. The people, supporting conditions or version of the programme may differ. A valid answer for the first population is not automatically the answer for the second.

External validity concerns whether an inference is justified beyond the particular people, settings, treatments, measurements and times directly studied. Evidence transfer makes that question concrete: what can this source study tell us about this specified target, and what assumptions are required to connect them? The task is neither to copy a result blindly nor to dismiss every setting as unique. It is to identify the differences that matter.

This guide uses original hypothetical education and library examples. Their numbers are constructed, not measured eduKate outcomes. The methodological literature includes research in health and education; citing a method does not transfer authority to make clinical recommendations or establish that an educational programme works.

Reading route: Begin with the target question, examine one programme with two population effects, then work through the assumptions behind transfer, implementation differences and the decision about what to do next.

“Does it work?” is missing most of the question

Before transferring evidence, write the intended claim in full. Which population? Which intervention or exposure? Compared with what? Which outcome? At what time after implementation? Delivered under which conditions? A sentence that omits these elements can make a narrow result sound universal without changing a single reported number.

For example, “the reading programme works” might refer to an immediate vocabulary assessment among volunteers receiving trained instruction. A decision-maker may actually need to know whether a different group will improve independent reading six months later when the programme is delivered within an already full timetable. The label is the same; the question is not.

Specify the source question and target question side by side. Some differences may prove irrelevant. Others may require new data, a different analysis or a narrower recommendation. The comparison is useful before any statistical model is fitted because it exposes what the model would have to bridge.

A target is not necessarily an entire country or every future learner. It can be a defined school intake, a set of library branches, an age group during a particular year or people meeting explicit eligibility criteria. Narrowing the target is not an evasion when that target is the one the decision actually concerns.

Internal validity and external validity protect different claims

Internal validity concerns whether the study supports its inference for the setting and population it actually addresses. External validity concerns the further step to another target. A study can fail at either stage. Making participants resemble a target population does not repair a biased treatment comparison; a strong treatment comparison does not, by itself, establish that every target will respond similarly.

Random treatment allocation and random sampling do different work. The first can support a fair comparison between assigned groups under the design’s conditions. The second concerns how observed units represent a wider population. A trial can randomise treatment among volunteers without randomly sampling everyone to whom a later recommendation will apply.

Elizabeth Tipton’s research on comparing experimental samples and inference populations addresses this problem in educational and behavioural research. Its similarity assessment is tied to selected observed covariates, not a guarantee that all relevant differences have been captured.

The practical implication is straightforward: inspect both the causal comparison and the selection into that comparison. Asking only whether a study was randomised leaves half of the transfer problem unexamined.

Generalisation and transportability: define the working distinction

Terminology varies across disciplines. In this article, generalisation means extending an inference to a broader population of which the study group is understood to be a part. Transportability emphasises moving an inference to a separately specified population or environment. The important requirement is to define the populations and assumptions, rather than win an argument over vocabulary.

Pearl and Bareinboim’s work on transportability across populations formalises how stated similarities and differences can determine whether experimental and observational information can be combined to identify a target causal effect. Their selection diagrams encode assumptions about differences; the diagram does not independently prove that those assumptions describe the world.

For a non-specialist reader, the transferable habit is to make the bridge explicit. “The two schools look alike” is not a bridge. “The programme’s effect is assumed to be comparable within these baseline groups, and we know how common those groups are in the target school” is a specific, challengeable proposal.

A useful bridge tells the next researcher where to disagree. It identifies the mechanism, the information needed and the point at which an unsupported extrapolation would begin.

A worked example: the same subgroup effects, different overall results

Imagine a fictional learning programme assessed in two baseline support groups. For simplicity, assume we know the group-specific mean outcomes under the programme and under its comparator. Treat these as constructed population quantities so that sampling error does not obscure the arithmetic.

Invented values used only to demonstrate population composition
Baseline support groupMean under comparatorMean under programmeProgramme effect
Higher support60688 points
Lower support50522 points

Suppose the source population is eighty per cent higher-support learners and twenty per cent lower-support learners. Its average programme effect is 0.8 × 8 + 0.2 × 2 = 6.8 points. Its comparator mean is 58 and its programme mean is 64.8.

Now suppose the target population has the reverse composition: twenty per cent higher-support and eighty per cent lower-support. Assume, for this example, that the same group-specific outcomes apply in the target. Its average effect is 0.2 × 8 + 0.8 × 2 = 3.2 points. Its comparator mean is 52 and its programme mean is 55.2.

No subgroup effect changed. No score was misreported. No institution necessarily failed. The populations contain different proportions of groups for whom the programme has different effects.

If a hypothetical implementation decision requires an average improvement of at least four points, the source result clears the threshold and the target result does not. The programme is not simply effective or ineffective in a context-free sense. The relevant effect belongs to a population, comparison and outcome scale.

This example illustrates the problem described in Dahabreh and colleagues’ tutorial on extending trial inferences to a new target population: differences in the distribution of treatment-effect modifiers can make a source average unsuitable as the target average.

An effect modifier is not merely a predictor of outcomes

In the constructed example, baseline support is an effect modifier because the programme-comparator difference is eight points in one group and two in the other. It also predicts outcome levels, but those are separate roles.

To see the distinction, change the fictional table so that both groups improve by five points. Their comparator means could still differ by ten points. Moving from one population mixture to another would change the overall outcome level but not the average programme effect on this additive scale, provided the subgroup effects really remained five everywhere.

That is why a table of demographic differences is not the end of an external-validity assessment. Some differences may strongly predict outcomes without modifying the effect being transferred. Others may matter only jointly. A variable’s relevance depends on the mechanism and the particular effect scale.

The analyst should not label every recorded characteristic an effect modifier simply because it is available. Nor should a plausible modifier be dismissed because a small study fails to detect an interaction. The first choice overcomplicates the bridge; the second can overlook a difference the study was not capable of estimating precisely.

Standardisation changes the mixture, not the evidence inside each group

The arithmetic above is standardisation in its simplest form. Estimate or otherwise justify an effect within each relevant group, then average those effects using the target population’s group proportions rather than the source proportions.

Target average effect = sum of [target group proportion × transferable group-specific effect].

This equation is an accounting identity once the quantities inside it are defined and justified. The difficult part is not multiplication. It is establishing that the group-specific effects can legitimately travel and that the target proportions are known well enough.

In our two-group example, the adjustment is visible. In a real dataset, relevant characteristics may be continuous, numerous and unevenly distributed. Models may then estimate outcome relationships or participation probabilities rather than divide people into a small table. Such modelling adds assumptions and estimation uncertainty; the simplicity of the conceptual formula does not remove them.

A useful report keeps the original source estimate and the target-adjusted estimate distinguishable. The latter is not a corrected version of a mistaken study result. It is an estimate of a different population quantity.

What has to hold for the bridge to carry weight?

At minimum, examine the source study’s own identification conditions, whether the relevant treatment and outcome are comparable, whether the selected baseline information captures the differences needed for the proposed transfer, and whether the source contains support for the target’s relevant groups.

Conditional exchangeability is a technical way of expressing an assumed comparability after accounting for specified characteristics. The precise version required depends on the target quantity and method. Transferring an average effect can sometimes require weaker assumptions than transferring both potential-outcome means separately. A methods specialist should match the assumption to the estimator rather than use the phrase as a universal permission slip.

Selection positivity or overlap means that target-relevant characteristics must have appropriate representation in the source for the intended method. A model cannot recover empirical evidence for a wholly absent group simply by assigning it a large weight.

These assumptions are central to the methods compared in the trial-to-target tutorial. They should appear in the interpretation, not remain hidden in an appendix while the headline presents an unconditional target claim.

Reweighting helps with imbalance, not with an empty region

Suppose the source includes some learners from both support groups but in the wrong proportions for the target. Reweighting can alter their influence so that the analysis better reflects the target mixture, under the stated conditions.

Now suppose the source contains no lower-support learners. Our target is eighty per cent lower-support. There is no observed source effect of two points to insert into the formula. Filling that position with the higher-support effect of eight points is an extrapolation, not a consequence of weighting.

The honest options include narrowing the target, obtaining evidence about the missing group, presenting sensitivity scenarios or acknowledging that the desired effect is not identified from the available material. Which option is useful depends on the decision and its consequences.

Near-absence matters too. A handful of records carrying most of the target weight may produce an unstable estimate. A smooth fitted curve can conceal that fragility. Inspect where evidence is plentiful, sparse and absent before interpreting a single target average.

This is also a useful library principle: a cross-reference to a broad owner does not mean that the owner contains evidence for every specialised use. Navigation can find a source; it cannot create coverage the source lacks.

A larger source study does not automatically solve a transfer problem

Imagine increasing the fictional source study a hundredfold while preserving its eighty-twenty composition and its restricted implementation conditions. Its average effect could become very precisely estimated at 6.8. The target average in our constructed world would still be 3.2.

Size improves the precision of particular estimates; it does not turn one population into another. Additional recruitment is most helpful for transfer when it supplies information about the groups, settings or mechanisms that the target claim depends on.

Design can address this earlier. Tipton’s research on stratified sampling for improved experimental generalisation examines selection strategies aimed at improving the relationship between experimental samples and inference populations. The practical lesson is to plan for the intended target before recruitment, rather than ask an unsuitable completed sample to answer every later question.

For a hypothetical library pilot, that could mean including branches with materially different visitor patterns and staffing conditions, not simply increasing observations within the easiest flagship branch. The proposed design would still need a clear purpose, feasible recruitment and an analysis appropriate to branch-level dependence.

The same programme name can conceal a different intervention

A programme can change when its duration, materials, support, delivery channel, staffing or participation requirements change. A trial of guided sessions does not automatically evaluate a self-directed version merely because the worksheet has the same title.

VanderWeele and Hernán’s work on multiple versions of treatment explains why treatment definitions and versions matter to causal interpretation. In this article’s practical examples, the corresponding task is to describe what people actually receive rather than rely on an umbrella label.

Consider a fictional library enquiry service. The source branch supplies appointments with specialist staff and access to a curated database. The target branch plans a walk-in desk with general staff and no database subscription. Calling both arrangements research support does not make their operative components equivalent.

An adaptation is not necessarily a mistake. It may be exactly what the new setting requires. But adaptation creates a new evidential question: which functions must remain intact, which can change, and how will the changed version be assessed? The original result should not silently become a guarantee for the adaptation.

The comparator can change even when the programme does not

Effects are contrasts. A programme compared with no additional support answers a different question from the same programme compared with a strong existing service. The word effective can conceal this difference.

In a constructed case, a programme produces a mean outcome of 75 in both settings. If the source comparator produces 65, the contrast is ten points. If the target comparator produces 72, the contrast is three points. The programme’s outcome did not deteriorate; the alternative improved.

For a real adoption decision, the relevant comparator is usually what people would otherwise receive in the target, not whichever alternative made the original study easiest to interpret. A new service can be valuable relative to one baseline and unnecessary relative to another.

Document the comparator with the same care as the intervention. Include normal support, access to substitutes, likely behaviour and the relevant time period. An intervention cannot be meaningfully transported while its contrast is left behind.

Measurement must travel with the claim

A result expressed in assessment points belongs to an assessment and an interpretation of those points. A target institution may use a different task, language, scoring practice or follow-up interval. Before transferring the number, establish whether it still refers to a comparable construct.

For example, a hypothetical vocabulary activity might improve recognition of recently practised words. A target decision may concern spontaneous use in unfamiliar writing. Both are related to vocabulary, but improvement on one task does not mathematically establish an equal improvement on the other.

The general research-methods distinction is developed in Research Methods and Source Evaluation. Here the application is specific: compare what was measured, how it was scored and when it was measured before carrying an effect size into a new setting.

Sometimes a link can be justified through additional validation evidence. Sometimes the original finding should be treated as evidence about a related mechanism rather than a direct estimate of the target outcome. Neither route should be disguised as direct observation of an outcome the study never collected.

Absolute and relative effects make different transport assumptions

Suppose a hypothetical intervention reduces an undesirable event probability from 0.20 to 0.10 in the source population. Its absolute reduction is 0.10; its risk ratio is 0.5. Those are two descriptions of the same source result, but assuming they remain constant in the target produces different predictions.

At a target baseline probability of 0.05, carrying over the ratio of 0.5 predicts a treated probability of 0.025. Carrying over the absolute reduction of 0.10 predicts −0.05, which is impossible for a probability. The algebra exposes that both scales cannot be held constant across all baseline risks.

This is a mathematical illustration, not an estimate of any health or educational intervention. It shows why the scale of the claimed effect belongs in the transfer argument. A convenient scale is not automatically a stable scale.

When decisions concern actual numbers of affected people, report target-relevant absolute quantities where justified, alongside the assumptions that produce them. A relative measure alone may not tell a decision-maker how much difference adoption would make in the target.

Time and scale can change the system being studied

A result can travel across geography and still fail to travel across time. The target population may have new tools, different prior experience or changed alternatives. A programme that addresses a scarce resource in one period may face a different constraint later.

Scale also introduces new questions. In a fictional pilot, a specialist team might personally support every participating branch. Extending the service to hundreds of branches changes the support capacity. If that support is part of the mechanism, the rollout does not simply repeat the pilot many times.

There may also be spillovers. One branch’s new service could redirect demand from another. Learners in different assigned groups could share materials. These interactions mean that outcomes may depend on how widely the intervention is used, not only on one person’s assignment.

These are prompts for a target-specific design, not claims that every rollout must fail. Write down which capacity constraints and interactions might change. Then decide whether existing evidence addresses them, whether a staged rollout could learn about them or whether the planned transfer is too large a leap.

Several studies help only when their differences remain visible

Evidence from multiple sites can reveal whether effects vary and whether a proposed transfer model behaves plausibly across settings. But pooling studies into one average does not automatically produce the effect in a new target population.

Dahabreh and colleagues’ work on causally interpretable meta-analysis addresses transporting information from multiple randomised trials to a target. Its focus on a specified population is different from treating a pooled research average as a universal answer.

For a library, an evidence table should therefore retain the study population, intervention version, comparator, outcome, setting and time. A summary that discards those fields may become shorter while losing the information needed to judge applicability.

Disagreement is not always a defect to average away. It may expose an effect modifier, a measurement difference or a change in implementation. It may also arise from sampling error or study bias. The useful task is to investigate the structure of disagreement rather than choose between unquestioning pooling and complete abandonment of synthesis.

An unmeasured difference should become a sensitivity question

Suppose the two-group transfer gives a target effect of 3.2 points, but we worry that target delivery conditions reduce the benefit further. We can represent a hypothetical additive decrement b and examine a target effect of 3.2 − b. Setting b to 0, 1 or 2 produces 3.2, 2.2 or 1.2 points. These are scenarios, not new empirical findings.

The value of the exercise is to ask what evidence would make a particular decrement plausible. A sensitivity range needs a substantive rationale. Assigning b = 0 because it preserves a preferred result is not an investigation.

Methods using bias functions and research published in June 2026 on global sensitivity for trial-to-target inference examine departures from the assumptions used in such transfers. The simple arithmetic here illustrates the purpose without claiming to implement those methods.

A useful conclusion identifies what would overturn the recommendation. The phrase context matters becomes operational when it names a possible difference, a consequence and a way to investigate it.

A pilot should answer the unresolved transfer question

“Run a pilot” is sensible only after specifying what the pilot needs to reveal. A small uncontrolled trial of delivery may teach an institution about staff workload, uptake or practical barriers. It may not estimate a causal learning effect with useful precision.

In the library example, uncertainty about appointment demand calls for different observation from uncertainty about the quality of research support. Testing whether staff can operate a booking system does not establish whether users receive better answers. Testing user satisfaction does not automatically establish factual accuracy.

Set the pilot’s questions, measures and decision rules before interpreting its outcome. Preserve the difference between feasibility, acceptability, effectiveness and cost. A pilot can successfully reveal a barrier even when it does not justify expansion.

A small pilot also cannot demonstrate the absence of rare problems merely because none occurred. Its information should be used at the resolution the design supports. Strong learning comes from a clear question and a fit-for-purpose observation, not from attaching the word pilot to any initial attempt.

The transfer decision is a choice among actions, not a verdict on all research

An institution may adopt, adapt, test, defer or decline. It may also adopt for one group while gathering evidence for another. These are different responses to evidence, not simply different levels of enthusiasm.

For our constructed target, an estimated effect of 3.2 below a four-point threshold could prompt redesign, a search for a less costly alternative or investigation of whether the threshold and payoff model are appropriate. It would not justify claiming that the source effect of 6.8 was false.

Conversely, a favourable transported estimate is not the entire decision. Feasibility, opportunity cost, rights, burden and other outcomes can matter. Some constraints are non-negotiable and should define admissible options before a benefit calculation begins.

A decision record should preserve the estimated benefit, its uncertainty, the transfer assumptions and the next review trigger. When a condition changes, the institution can revisit the relevant part of the argument rather than either defending the old conclusion indefinitely or restarting from nothing.

An evidence-transfer brief a reader can actually use

Begin with one paragraph stating the target population, intervention version, comparator, outcome and time horizon. Follow it with a compact account of the source evidence and why it supports its own claim. Then identify the similarities and differences that are relevant to the proposed bridge.

State which quantities come directly from the source, which come from target data and which are assumptions. Show the target calculation or reasoning. Describe regions of weak overlap, unresolved implementation changes and sensitivity to unmeasured differences.

End with the decision the evidence supports and the next observation that could change it. A useful sentence is: “This evidence supports a conditional estimate for the specified target, assuming comparable subgroup effects and delivery; the target group absent from the source remains unresolved.”

A weak sentence is: “Research proves that this works for everyone.” It may be easier to market, but it gives neither a careful reader nor a future researcher a way to inspect the leap.

What this adds to a knowledge library

A library can contain excellent articles and still support poor decisions if a reader cannot tell where each claim applies. Source quality and applicability are different questions. A highly credible study can be only indirectly relevant to a particular target.

The solution is not to bury every article under warnings. It is to preserve the few boundaries that determine responsible use: population, intervention, comparator, outcome, time, method and uncertainty. Those fields allow an explanation to travel with its conditions attached.

For a learner, the habit is equally valuable. When encountering a persuasive example, ask: what is similar to my problem, what is different, and why should the difference matter? That question avoids both naïve imitation and the claim that nothing can ever be learned from somewhere else.

Evidence travels well when the bridge is visible. It travels badly when a number is detached from the conditions that made it meaningful.

Sources, boundaries and further reading

Source records and accessible methodological descriptions were checked on 5 September 2026. This is an explanatory synthesis, not a systematic review or a validated transport analysis. The worked calculations are original hypothetical examples. Applying technical estimators requires their full assumptions, suitable data and method-specific expertise.

Continue through the Library: Read Experimental Design for the source comparison, Causal Inference for causal identification, Comparative Systems Research for contextual comparison, and Systematic Reviews and Evidence Synthesis for combining studies. Return to the Research Collections Directory for the wider collection.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading