An experiment is not defined by having a laboratory, a white coat or a complicated instrument. It is defined by a deliberate change introduced so that the consequences of that change can be compared with what would otherwise have happened.
That sounds simple until reality interferes. People differ. machines drift. weather changes. classrooms have different teachers. animals have different histories. materials come from different batches. measurements contain error. participants know what treatment they received. researchers have expectations. some units disappear before the study ends.
Experimental design is the architecture that separates the effect we care about from these competing sources of variation. It asks, before data collection begins, how the experiment should be arranged so that the eventual comparison can support a credible causal inference.
The experimental design loop
QUESTION → CAUSAL CLAIM → FACTORS + RESPONSES → EXPERIMENTAL UNITS → CONTROL / COMPARATOR → RANDOMISE → BLOCK WHAT MATTERS → REPLICATE → BLIND WHERE POSSIBLE → MEASURE → CHECK PROTOCOL → ANALYSE → TEST ASSUMPTIONS → INTERPRET EFFECT + UNCERTAINTY → REPLICATE AGAIN
The central idea is that design comes before analysis. Sophisticated statistics cannot fully repair an experiment that confounded treatment with time, assigned all strong participants to one condition, changed the measuring instrument halfway through or collected only one observation per treatment.
1. Start with the causal question
“Does method A produce better outcomes than method B under these conditions?” is an experimental question. “Are people who use method A different from people who use method B?” is observational unless the researcher controls assignment.
A causal question should name the intervention or factor, the comparison, the response and the population or system to which the conclusion is intended to apply.
This connects directly to How Research Methods and Source Evaluation Work: the method must fit the kind of claim being made.
2. Define factors, levels and responses
A factor is a variable deliberately manipulated or studied. A factor’s levels are the conditions being compared. The response is the outcome measured after or during exposure to those conditions.
In a manufacturing experiment, temperature and pressure might be factors while defect rate is the response. In an education experiment, feedback type might be the factor while later retrieval performance is the response. In agriculture, fertiliser dose might be the factor and crop yield the response.
Precise definitions matter because vague factors and unstable responses create ambiguous experiments even when the mathematics is correct.
3. The experimental unit is the unit that receives the treatment
A common design error is to confuse the measurement unit with the experimental unit. If an entire classroom receives one teaching method, the classroom may be the experimental unit even if scores are collected from thirty individual students. Treating all thirty students as thirty independently randomised units can create pseudoreplication.
The experimental unit is determined by how the treatment is assigned, not by how many measurements are recorded afterward.
4. Controls create the comparison needed for causality
A control or comparator represents what the experimental units experience without the treatment of interest, or under an alternative treatment. The correct comparator depends on the question.
- No-treatment control: asks what happens without intervention.
- Placebo control: helps separate the intervention’s specific effect from expectations or treatment context where appropriate.
- Active control: compares against an existing intervention.
- Baseline or within-unit comparison: compares a unit with itself over time, when design assumptions support it.
- Sham procedure: may be used in some procedural research but carries ethical considerations.
A poorly chosen control can make an experiment answer a different question from the one readers think it answers.
5. Randomisation protects against systematic assignment bias
Randomisation assigns treatments by a random mechanism rather than by researcher preference, participant choice or convenient order. The NIST/SEMATECH Engineering Statistics Handbook describes completely randomised designs as designs in which factor levels are randomly assigned to experimental units.
Randomisation does not guarantee that every characteristic will be perfectly balanced in a finite experiment. Its deeper value is that treatment assignment is not systematically tied to known or unknown characteristics through investigator choice.
It creates a defensible basis for probabilistic inference about treatment effects under the design.
6. Random sequence is different from random sampling
Random assignment concerns which treatment an experimental unit receives. Random sampling concerns how units were selected from a larger population. An experiment can randomise treatment well while using a convenience sample.
Random assignment strengthens internal causal inference. Random sampling, when feasible, strengthens the basis for generalising to a target population. They solve different problems.
7. Replication estimates variation
Replication means applying the same experimental condition to more than one independent experimental unit. It allows researchers to observe how outcomes vary even when the treatment condition is nominally the same.
NIST’s design guidance emphasises replication because it provides information about process variation and supports tests of model adequacy. Without replication, a difference between treatments may be inseparable from an idiosyncratic difference between the particular units used.
Repeated measurements on the same unit are useful but are not automatically independent replication.
8. Blocking controls important nuisance variation
A nuisance factor affects the response but is not the main factor of interest. If it is known and important, the experiment can group similar units into blocks and compare treatments within those blocks.
NIST summarises the principle memorably: block what you can, randomise what you cannot. Blocking is useful when predictable sources of variation—such as batch, day, laboratory, classroom or location—could otherwise obscure the treatment effect.
Blocking is not the same as controlling away every difference. It is a design strategy for handling a small number of important nuisance factors explicitly.
9. Confounding is the enemy of interpretation
Two factors are confounded when their effects cannot be separated because they vary together in the design. If every treatment-A measurement is taken in the morning and every treatment-B measurement in the afternoon, treatment and time of day are confounded.
No amount of confidence in the treatment effect can remove the logical ambiguity: perhaps the difference was caused by treatment, time, or both.
Randomisation, blocking and balanced design are ways of preventing or managing confounding before it becomes embedded in the data.
10. Balance improves comparability
Balanced designs assign similar numbers of experimental units to the treatment combinations. Balance often simplifies analysis and improves efficiency because information is distributed evenly across conditions.
Real experiments can become unbalanced through dropout, missing data or practical constraints. Unbalanced designs are not automatically invalid, but analysis and interpretation become more dependent on modelling assumptions.
11. Factorial designs study more than one factor at once
A factorial experiment varies two or more factors so that main effects and interactions can be estimated efficiently. For example, an education study might vary feedback type and practice spacing. A manufacturing study might vary temperature, pressure and material formulation.
Factorial designs are powerful because real systems are interactive. The effect of one factor may depend on the level of another.
12. Interactions reveal conditional effects
An interaction occurs when the effect of one factor changes depending on another factor. If a study method helps novice learners but not experts, expertise may interact with method. If a chemical process responds to temperature only at one pressure level, pressure and temperature interact.
Ignoring interactions can make an average effect misleading. The system may not have one universal response to the factor.
13. Fractional factorial designs trade completeness for efficiency
When many factors are possible, testing every combination can become expensive. Fractional factorial designs run a strategically selected subset of combinations so that important effects can be estimated under assumptions about which interactions are negligible.
The trade-off is aliasing: some effects are deliberately confounded with others. Good design makes that confounding structure known rather than accidental.
14. Sequential experimentation can learn in stages
Not every experiment must be planned as one irreversible block. Researchers may begin with screening experiments to identify influential factors, then conduct narrower experiments to estimate interactions or optimise response.
Sequential design is useful when each experiment changes what the next question should be. The process remains rigorous when decision rules and new hypotheses are clearly distinguished from analyses planned in advance.
15. Blinding reduces expectation-driven bias
When participants, treatment providers, outcome assessors or analysts know which condition was assigned, expectations can influence behaviour, measurement or interpretation.
Blinding is not possible in every experiment. A student usually knows which teaching method they experienced. An engineer may know which material is being machined. The design should therefore identify who can reasonably be blinded and where knowledge of assignment creates the greatest risk.
“Double blind” is less informative than stating exactly who was unaware of which information.
16. Placebos separate treatment context from specific treatment effect
In some clinical experiments, placebo controls help account for expectations and treatment context. But placebo design is not universally appropriate, and ethical standards may require active controls when effective treatment exists.
The broader design lesson is that comparators should isolate the mechanism the experiment is trying to estimate without exposing participants to unjustified risk.
17. Measurement quality belongs inside experimental design
An experiment can assign treatments perfectly and still fail if the response measure is unreliable, insensitive or invalid. Measurement should therefore be designed alongside treatment assignment.
Researchers should define instruments, calibration, timing, units, scoring rules, assessor training and handling of repeated measurements before data collection where possible.
See How Scientific Measurement Works for traceability, calibration and uncertainty.
18. Outcome choice can create invisible bias
If researchers measure many outcomes but report only the ones that look favourable, the published experiment becomes a selected representation of the data. Pre-specifying primary and secondary outcomes helps distinguish planned inference from exploration.
Outcome switching is especially concerning when the choice is made after seeing results. Transparency about all measured outcomes is part of valid inference.
19. Sample size is a design decision, not an afterthought
Too small an experiment may estimate effects so imprecisely that important differences cannot be distinguished from noise. An unnecessarily large experiment can waste resources or expose more participants or materials than needed.
Sample-size planning depends on the design, expected variability, effect size of practical importance, acceptable error rates and desired precision or power. The calculation is only as meaningful as the assumptions supplied to it.
20. Statistical power is conditional, not magical
Power is the probability, under a specified alternative and analysis plan, that a statistical procedure will reject the null hypothesis. It depends on effect size, variability, sample size, significance threshold and design.
A high-powered experiment can still answer the wrong question or measure the wrong outcome precisely. Power is one property of a design, not a universal certificate of quality.
21. Preregistration separates confirmation from exploration
Preregistration records hypotheses, outcomes and analysis plans before outcomes are known. It does not forbid exploratory analysis. It helps readers distinguish what was predicted from what was discovered after examining the data.
Exploration is essential to science. The problem arises when exploratory findings are presented as though they were pre-specified confirmatory tests.
22. Protocol deviations should be visible
Real experiments encounter broken instruments, missing participants, unexpected safety issues and procedural errors. Deviations from protocol should be recorded because they may affect interpretation.
A deviation is not automatically fatal. Hidden deviations are more dangerous because later readers cannot judge their consequence.
23. Missing data can break randomisation after the fact
Random assignment balances treatment assignment at the start. If dropout differs systematically between groups, the analysed sample can become selected.
Researchers should track why data are missing, whether missingness differs by treatment and how analytical methods handle it. Complete-case analysis is not automatically unbiased.
24. Intention-to-treat preserves the logic of assignment
In randomised trials, intention-to-treat analysis generally compares participants according to the group to which they were assigned, even if adherence is imperfect. This preserves the original randomised comparison more directly than reclassifying participants by what they eventually did.
Other estimands may also be scientifically relevant, but they answer different causal questions and can require stronger assumptions.
25. Laboratory control and real-world relevance trade against each other
Highly controlled experiments can isolate mechanisms precisely, but artificial conditions may limit generalisation. Field experiments preserve realistic environments but face more uncontrolled variation.
The design should match the intended inference. Mechanism testing and operational effectiveness may require different experiments.
26. Internal validity and external validity are distinct
Internal validity concerns whether the experiment credibly estimates the effect within the studied setting. External validity concerns how well that effect transfers to other people, places, times or implementations.
Randomisation strengthens internal validity but does not guarantee broad external validity. A perfectly randomised study in one narrow population may not generalise beyond it.
27. Multi-site experiments test transportability
Running a common protocol across several schools, laboratories, hospitals or sites can reveal whether treatment effects depend on context. Site variation is not merely nuisance; it can identify conditions under which the intervention works differently.
Hierarchical designs and analyses may be needed because observations inside the same site are often correlated.
28. Cluster randomisation changes the unit of inference
Sometimes entire schools, villages, wards or work teams are randomised because individual assignment is impossible or contamination would occur. Outcomes may still be measured on individuals, but the clustering must be reflected in design and analysis.
Ignoring within-cluster correlation makes the experiment look more precise than it is.
29. Cross-over designs use participants as their own controls
In a cross-over design, units receive multiple treatments in different periods. This can be efficient when treatment effects are reversible and carryover can be managed.
Cross-over designs are inappropriate when treatment permanently changes the system, disease state evolves quickly or period effects make comparisons unstable.
30. Adaptive experiments modify parts of the design using accumulating data
Some experiments pre-specify rules for changing allocation, dropping ineffective arms, changing sample size or stopping early. Adaptive design can improve efficiency, but adaptation must be built into the inferential framework so that error rates and interpretation remain valid.
Changing the experiment informally whenever results look interesting is not an adaptive design; it is uncontrolled optional stopping.
31. Stopping rules belong in the design
Experiments may stop early for efficacy, harm, futility or operational reasons. Repeatedly checking results and stopping as soon as a threshold is crossed can inflate false-positive risk unless the analysis accounts for the monitoring plan.
Ethical and statistical stopping rules should therefore be planned together.
32. Negative results are informative when the experiment was capable of detecting something meaningful
“No statistically significant difference” does not automatically prove treatments are equivalent. The experiment may have been too imprecise.
Confidence intervals around the effect estimate help show which effect sizes remain compatible with the data. Equivalence and non-inferiority questions require designs and margins specifically built for those claims.
33. Analysis should follow the randomisation structure
The statistical model should respect how units were assigned, blocked, clustered and repeatedly measured. If the design contains paired units, clusters or blocks, the analysis should not pretend observations arose from a simple independent sample.
Design and analysis are one system. Good experiments make the eventual analysis possible rather than handing statisticians an accidental data structure after collection ends.
34. Graphs and raw data can reveal failures that summary tests hide
Plotting responses by treatment, time, batch or site can reveal drift, outliers, nonlinear relationships, variance changes or data-entry errors. A single p-value compresses all of this structure.
Exploratory visualisation should accompany formal analysis, with the distinction between exploration and pre-specified testing kept visible.
35. Replication across experiments is stronger than repetition inside one experiment
An experiment may contain many internal replicates and still depend on one laboratory, one operator, one instrument or one moment in time. Independent replication tests whether the finding survives a new implementation.
Conceptual replication goes further by testing the same underlying claim using a different operationalisation or method. Converging evidence across designs can strengthen the mechanism beyond one procedural recipe.
36. Experiments become evidence only after reporting
A result cannot be evaluated if the report omits assignment method, sample size, exclusions, protocol deviations, outcome definitions or analysis choices. Transparent reporting is therefore part of the experimental evidence chain.
The later systematic review depends on these details. Poor reporting weakens future synthesis even when the original experiment was well designed.
37. Experimental design in the eduKate Library
This article owns the general architecture of controlled experimentation. It does not take ownership from Medicine’s clinical-trial pages, Biology, Veterinary science or domain-specific laboratory articles. Those domains own the substantive safety, ethics and interpretation of experiments in their fields.
The general method routes outward to How Scientific Research Works, How Laboratory Practices Work and How Systematic Reviews and Evidence Synthesis Work.
38. A practical experiment-design audit
- What causal question is being asked?
- What is the experimental unit?
- What factors and levels are manipulated?
- What response is measured?
- What comparator answers the intended question?
- How is treatment assigned?
- What nuisance factors should be blocked?
- Is there true independent replication?
- Could treatment be confounded with order, site or time?
- Who can be blinded?
- Is the measurement valid and reliable?
- Was sample size planned for practical precision or power?
- Are outcomes and analysis plans pre-specified?
- How will missing data and deviations be handled?
- Does the analysis match the randomisation structure?
- What population and settings can the result reasonably generalise to?
39. Experimental design is organised doubt
An experiment is valuable because it makes a counterfactual comparison more credible. Randomisation says we did not choose who received what. Blocking says we recognised an important nuisance source. Replication says we did not trust one unit. Blinding says we anticipated expectation. Preregistration says we distinguished prediction from discovery.
The design is therefore a visible record of what the researcher did not want to leave to faith.
Sources and authoritative guidance
- NIST/SEMATECH Engineering Statistics Handbook
- NIST — Completely Randomized Designs
- NIST — Randomized Block Designs
- NIST — Blocking of Factorial Designs
- NIH — Rigor and Reproducibility
Continue through eduKate
- How Research Methods and Source Evaluation Work
- How Scientific Research Works
- How Laboratory Practices Work
- How Scientific Measurement Works
- How Systematic Reviews and Evidence Synthesis Work
- Data Quality
Wintour House return: A good experiment does not ask the reader to trust the researcher’s intention. It arranges the comparison so that alternative explanations are weakened before the result is known, preserves the assignment and measurement logic, and leaves enough detail for another team to test whether the effect survives a new encounter with the world.