How Clustered and Multilevel Data Work | When Observations Belong to Groups, Places and Repeated Contexts

Two hundred learners complete the same assessment. A spreadsheet therefore contains two hundred rows. It is tempting to treat those rows as two hundred independent pieces of evidence.

But the learners belong to ten classes. They share teachers, timetables, peer environments, lesson sequences and sometimes the same disruptions. Two learners from the same class may therefore resemble one another more than two learners chosen from different classes.

Clustered data arise when observations share a higher-level context that can make them statistically dependent. Multilevel data make those levels explicit so variation can be studied at the level of the observation, person, group, place, organisation or time structure where it occurs. The key problem is not that groups exist. It is that pretending dependence does not exist can produce confidence that is too strong, standard errors that are too small, or interpretations that confuse individual and group-level processes.

This article builds from classes and learners to repeated measurements, organisations, cross-classified systems and cluster-randomised designs. The numerical examples are constructed for explanation; they are not eduKate class data, outcome claims or programme evaluations.

Reading route: begin with why independence fails, calculate the intraclass correlation and design effect, then move through multilevel models, partial pooling, cluster assignment, cross-classification and the reporting contract.

Rows can be separate without being independent

Statistical independence is stronger than physical separateness. Two essays can be written by different learners and still share influences because both learners experienced the same teacher, task briefing and classroom environment.

The same structure appears almost everywhere: customers within branches, repeated readings within sensors, workers within teams, residents within neighbourhoods, transactions within accounts, pages within websites and measurements within the same person over time.

When observations share a context, some information is duplicated at the group level. Twenty learners in one class give much more information about individual variation than one learner, but they do not provide twenty independent replications of the teacher or classroom environment.

The first task is therefore structural: draw the data-generating hierarchy before choosing the statistical model.

The hierarchy belongs to the research question

A dataset can contain several valid hierarchies. Assessment occasions may be nested within learners, learners within classes, classes within schools and schools within districts. A question about short-term learner change may require repeated occasions within learners. A question about school policy may need the school level explicitly represented.

Not every available level must appear in every model. Include the levels needed to represent the sampling, assignment, dependence and mechanism relevant to the claim.

Conversely, flattening the hierarchy because a CSV file has one row per observation does not remove the structure. File format is not research design.

Intraclass correlation measures resemblance inside a cluster

For a simple random-intercept setting, the intraclass correlation coefficient, ICC, can be written as the between-cluster variance divided by the total variance:

ICC = between-cluster variance ÷ (between-cluster variance + within-cluster variance).

An ICC of zero corresponds to no extra resemblance from sharing a cluster in that simple model. An ICC near one means observations within a cluster are very similar relative to observations from different clusters.

The ICC is outcome- and context-specific. The same classes can have a high ICC for one behaviour and a low ICC for another. It should not be treated as a permanent property of a school or organisation.

A small ICC can matter when clusters are large

For equal cluster size m in a simple setting, a familiar approximation to the design effect is:

Design effect = 1 + (m − 1) × ICC.

Suppose ten fictional classes each contain twenty learners, giving 200 learner records, and the ICC for the outcome is 0.10. The design effect is 1 + 19 × 0.10 = 2.9.

A rough effective-sample-size intuition is 200 ÷ 2.9, about 69 independent observations. This is not a universal replacement formula for every estimator or unequal-cluster design. It demonstrates why a modest within-class correlation can substantially reduce the amount of independent information when many observations share each class.

The 2012 CONSORT extension for cluster randomised trials highlights the need to account for clustering in sample-size calculations and analysis and to report measures such as intracluster correlation where relevant.

Ignoring clustering usually makes uncertainty look too small

If an ordinary analysis assumes every learner is independent when learners within classes are positively correlated, the standard error can be underestimated. Confidence intervals become too narrow and conventional tests can reject too often.

The point estimate itself may or may not change substantially depending on the model and design. The most immediate failure is often confidence: the analysis behaves as though it observed more independent contexts than it actually did.

This is a recurring research principle. More rows are not always more independent evidence. Independence is a property of the information-generating structure, not the spreadsheet length.

The unit of assignment can be different from the unit measured

In a cluster-randomised study, entire schools, classes, villages or clinics may be assigned to conditions while outcomes are measured on individuals. The number of individuals can be large while the number of independent assignment units is small.

Twenty schools with fifty learners each do not create one thousand independently randomised schools. The treatment contrast is anchored in twenty cluster assignments.

Analysis should respect the design. The CONSORT 2025 and SPIRIT 2025 ecosystem continues to distinguish trial designs and their specialised extensions, including clustered designs.

Cluster randomisation changes recruitment and interpretation

When clusters are assigned before individuals are recruited, knowledge of the cluster’s assignment can influence who enters the study. This creates selection pathways that do not arise in the same way when individuals are randomised after recruitment.

Cluster-level interventions can also spill across individuals by design: a teacher changes classroom practice for everyone; a branch changes its service workflow; a neighbourhood receives an infrastructure change.

The estimand must therefore say whose outcome is being compared under which cluster-level assignment and handling of events after assignment. A recent CRT-Estimands Framework extends the ICH E9(R1) estimand logic specifically for cluster-randomised trials and illustrates how the research question and analysis should be aligned.

A multilevel model lets variation exist at several levels

A simple two-level model can represent an outcome as a grand average, effects of measured predictors, a cluster-specific deviation and an individual-level residual. The cluster deviation acknowledges that some classes or branches sit systematically above or below the overall level.

This allows the analysis to answer several questions at once: how much variation lies between clusters, how much lies within clusters, whether a predictor operates within or between clusters, and how uncertainty changes when group structure is recognised.

The phrase “random effect” describes a statistical component, not randomness in the everyday sense. The model treats a set of cluster deviations as arising from a distribution whose variance is estimated. Whether that is appropriate depends on the inferential target and design.

Fixed effects and random effects answer different structural questions

If the analysis cares only about a small named set of classes and wants a separate coefficient for each, fixed class indicators may be appropriate. If the classes are viewed as a sample from a wider population of possible classes and the goal includes estimating between-class variation, a random-effects structure can be useful.

The choice is not merely computational. It changes what is estimated and how information is shared. Hybrid models are common: some group-level features are represented explicitly while unexplained group variation remains stochastic.

Do not choose a random effect only because software defaults to one or a fixed effect only because the number of groups is small. Begin with the target inference and the assignment or sampling process.

Partial pooling balances local evidence with the wider pattern

Suppose one small class has an observed mean much higher than every other class. A no-pooling analysis reports its mean exactly as observed. A complete-pooling analysis ignores class identity and gives every class the same estimated mean.

A multilevel model can partially pool: the class estimate is informed by its own data and by the distribution of class effects. Small, noisy classes are generally pulled more toward the overall pattern than large, precisely observed classes.

This shrinkage is not falsification of the observed mean. It is an estimate under a hierarchical model that recognises uncertainty in small-group extremes. The raw group mean and model-based estimate answer different questions and should remain distinguishable.

Partial pooling is not always appropriate

Pooling assumes some exchangeable structure among groups after accounting for modeled differences. A one-off specialised institution may not belong to the same population as ordinary branches. A cluster created by a unique policy regime may deserve explicit treatment rather than automatic shrinkage toward the rest.

Inspect the grouping mechanism. If groups differ in known structural ways, represent those differences. Hierarchical modelling is powerful because it borrows information; that borrowing must have a defensible source.

Within-group and between-group relationships can differ

Suppose classes with more study time on average also have higher scores. That between-class relationship does not establish that a learner who studies one additional hour within the same class will gain the between-class amount.

The class average can capture teacher expectations, timetable structure, prior attainment or other contextual factors. A predictor that varies both within and between groups can therefore carry two distinct associations.

Multilevel analysis can separate an individual’s deviation from their group mean from the group mean itself. This prevents one coefficient from silently mixing within-group and between-group information.

Centering is a meaning choice, not cosmetic preprocessing

Subtracting a group mean from an individual predictor creates a within-group deviation. Subtracting the overall mean changes the interpretation of intercepts while preserving different information.

There is no universal rule that every multilevel predictor must be group-mean centred. The correct transformation depends on the estimand and whether within- and between-group effects need to be separated.

Document centering choices. A coefficient can change meaning even when the underlying data values have merely been re-expressed.

The ecological fallacy runs from groups to individuals

If districts with more libraries have higher literacy, it does not follow that the individuals who use libraries are the individuals with higher literacy, or that adding a library would necessarily cause the district relationship to appear within people.

Group-level associations can arise from composition, context or confounding. An ecological analysis belongs to the group level unless a justified bridge to individuals is supplied.

The reverse mistake is sometimes called the atomistic fallacy: an individual-level relationship need not reproduce at the group level. Multilevel thinking protects both directions by keeping the level of each claim visible.

Contextual effects are not just aggregated individual effects

A learner’s own reading time may matter, and the average reading culture of the class may matter beyond that individual’s reading time. Those are different mechanisms.

A contextual effect compares individuals with similar personal characteristics who belong to groups with different group-level conditions. It requires stronger interpretation than merely noticing that group averages differ.

As always, association is not automatically causation. Multilevel structure helps represent where variables live; it does not remove confounding or selection bias.

Repeated measures create clustering within people

Measure the same learner every month and the observations share a person. Treating twelve monthly scores as twelve independent learners would exaggerate information.

A longitudinal multilevel model can represent a learner-specific baseline and trajectory, while residual correlation can capture additional time structure. Time can be continuous, categorical or piecewise depending on the mechanism.

The structure also clarifies missing data. A learner with one missing month is different from an entire class that leaves the study. Missingness can operate at several levels; route those questions to Missing Data Analysis.

Serial correlation and clustering are related but not identical

Repeated observations close in time can be more similar than observations far apart even after accounting for a person-specific effect. That is a correlation structure within the lower-level residuals.

A random intercept captures persistent person-level similarity. An autoregressive residual can capture extra temporal dependence. One does not automatically replace the other.

Use the structure required by the data-generating process rather than adding every possible correlation term. Overly complex models can become unstable, especially with few clusters or short series.

Not every hierarchy is neatly nested

A learner may attend one school but be taught mathematics by one teacher and English by another. A research paper can be reviewed by several reviewers who each review papers from many journals. A patient can receive care from several clinicians.

These structures are cross-classified rather than strictly nested. Forcing them into one simple tree can attribute variation to the wrong level.

Multiple-membership models go further when an observation is influenced by several higher-level units at once. The correct structure follows the actual exposure or membership graph.

Unequal cluster sizes change precision and sometimes the estimand

Real classes and branches rarely contain exactly the same number of observations. A few very large clusters can dominate an individual-weighted analysis while a cluster-weighted analysis gives each cluster equal influence.

Neither weighting rule is universally correct. They target different averages. If the question is the average outcome for an individual, larger clusters naturally contain more individuals. If the question is the average cluster, equal cluster weighting may fit better.

State the target. A disagreement about weights can be an estimand disagreement rather than a technical dispute.

Informative cluster size creates another dependency

Suppose high-performing classes tend to remain small while struggling classes are merged into larger groups. Cluster size is now related to the outcome process.

Methods that assume cluster size is unrelated to outcomes may target a different quantity or become biased. Investigate why clusters have the sizes they do rather than treating size as a harmless administrative field.

The same issue appears in organisations where large branches differ systematically from small branches. Size can be a mechanism, a confounder or part of the target definition.

Cluster-robust standard errors and GEE offer population-average alternatives

Not every clustered problem requires a multilevel random-effects model. Cluster-robust standard errors can adjust uncertainty for within-cluster dependence under suitable conditions. Generalised estimating equations can estimate population-average associations with a working correlation structure.

These approaches and multilevel models can answer different versions of the question. A subject-specific logistic regression coefficient from a random-effects model is not generally identical in interpretation to a population-average GEE coefficient.

Select the estimator from the target effect, cluster count, dependence structure and robustness needs. Do not treat methods as interchangeable because they all accept a cluster identifier.

A small number of clusters is a serious information limit

Many asymptotic methods behave best with a reasonably large number of independent clusters. Hundreds of individuals inside six schools do not create hundreds of cluster-level replications.

With few clusters, variance estimates can be unstable and standard cluster-robust approximations can perform poorly. Small-sample corrections, randomisation-based methods or cluster-level analyses may be more appropriate in particular designs.

The key is not to hide the limitation behind a large individual N. Report both the number of clusters and the number of lower-level observations.

Cluster-level confounding cannot be repaired by many individuals

Suppose only two schools adopt a new programme and both are unusually well resourced. Measuring ten thousand learners in those schools does not create a strong causal comparison with ordinary schools. The programme and the school context remain entangled.

More observations within a confounded cluster can estimate that cluster’s outcomes precisely while leaving the between-cluster causal question unresolved.

This connects to Causal Inference: clustering is a dependence problem, while confounding is an identification problem. One analysis may need to address both.

Cross-level interactions can be scientifically meaningful

An individual-level relationship can differ depending on group context. The association between personal study time and achievement might be stronger in classes with high-quality feedback, for example.

A cross-level interaction represents that possibility explicitly. It should be motivated by a mechanism rather than discovered from a large search over every individual and group predictor.

Interactions can require much more data than main effects, especially at the higher level. The effective information for a class-level moderator depends heavily on the number and diversity of classes, not only the number of learners.

Multilevel models can overfit too

Random slopes, cross-level interactions, nonlinear time effects and several nested levels can create a model whose complexity exceeds the information in the data. Convergence warnings, boundary estimates and implausibly large correlations are signals to investigate, not messages to suppress.

Begin with the structure necessary for the question. Add complexity when it represents a real mechanism or materially improves fit and inference. Use sensitivity analysis to test whether the substantive conclusion depends on fragile covariance assumptions.

See Sensitivity Analysis and Robustness Checks for the broader discipline of exposing assumption dependence.

Prediction at a new cluster is different from prediction inside a known cluster

If a model has learned that Class 7 tends to score above average, predictions for another learner in Class 7 can use that class information. Predicting for a completely new class cannot use a class-specific effect that has not been observed.

Training and test splits should respect this target. Randomly splitting learners from the same class across training and test sets evaluates within-known-class prediction. Holding out entire classes evaluates generalisation to new classes.

The second task is harder and often closer to deployment in a new institution. Model validation should match the level at which the system will encounter novelty.

Multilevel structure matters for AI datasets too

A machine-learning dataset can contain millions of records but only a small number of sources, speakers, institutions or devices. Random row-level train/test splits can then leak source-specific patterns across both sides.

Group-aware splitting is one way to test whether a model transfers across the higher-level units that matter. If deployment requires performance on unseen schools, writers, hospitals or devices, validation should include unseen higher-level units where feasible.

This is not a special rule for AI. It is the same multilevel principle expressed in a predictive workflow.

Visualise the levels before trusting the coefficient

Plot group means and individual observations. Show trajectories for repeated measures. Display cluster sizes. Examine whether residuals or effects vary systematically across groups.

A single regression line can hide a strong between-group gradient and a weak within-group relationship. Small multiples can reveal whether one group drives the entire pattern.

Visualisation does not replace modelling, but it makes the hierarchy inspectable. Route chart design through Data Visualisation.

A reporting contract for clustered data

Report every relevant level, the number of units at each level and the typical and range of cluster sizes. State how clusters were formed, sampled or assigned. Explain whether the analysis targets individuals, clusters or another weighted population.

Describe how clustering entered sample-size planning and analysis. Report ICCs or other dependence measures where they are meaningful, with enough context for readers to interpret them.

For multilevel models, identify fixed and random components, covariance structure, centering decisions and estimation approach. For GEE or cluster-robust methods, report the clustering unit, working correlation or correction used where relevant.

For cluster trials, report the flow of both clusters and individuals and align the estimand, design and analysis. Current reporting resources can be found through the SPIRIT–CONSORT website and its design extensions.

A learner’s hierarchy check

  1. What is one row?
  2. Which rows share a person, class, place, device or organisation?
  3. At which level was treatment or exposure assigned?
  4. At which level does the claim live?
  5. How many genuinely independent higher-level units exist?
  6. Could within-group and between-group relationships differ?
  7. Will the model be used in known groups or new groups?

Those questions often diagnose the problem before any advanced statistics are required.

Where this article sits in the eduKate Library

This article owns the general structure of dependent observations across nested, repeated and grouped data. It does not replace education, medicine, biology, veterinary or organisational domain owners. Those domains determine which outcomes and mechanisms matter; this article provides the common statistical geometry underneath them.

Use Surveys and Sampling for how units enter a study, Experimental Design for assignment, External Validity and Evidence Transfer for movement to new populations, and Statistical Inference and Uncertainty for inferential foundations.

Sources and further reading

Source pages were checked for this edition on 5 September 2026. The ten-class design-effect calculation is an original illustration. Clinical-trial reporting documents are used here for their unusually explicit treatment of clustered design; they do not make this a clinical guidance article.

Continue through the Library: Surveys and SamplingExperimental DesignExternal Validity and Evidence TransferResearch Collections Directory.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading