A survey interviews 1,000 people. Another survey interviews 10,000. Which one represents the population better?
You cannot answer from sample size alone. You need to know who had a chance of selection, who actually responded, how those chances differed, what parts of the population were missed, and how the final estimates account for those differences.
Survey weighting is the controlled translation from observed sampled units to a target population. A weight tells an estimator how much population representation a responding unit carries under a stated sampling and adjustment procedure. Calibration and nonresponse adjustment can improve that translation by using selection probabilities and reliable auxiliary information. They cannot create direct evidence for population groups that the survey never reached.
This article owns the detailed weighting problem. The broader owner remains How Surveys and Sampling Work; incomplete outcomes are treated in How Missing Data Analysis Works. The examples below are fictional and are not eduKate operational data.
The simplest weight begins with selection probability
If every member of a population has the same chance of selection, an unweighted sample mean can estimate the population mean under the design. But many surveys deliberately give different units different selection probabilities.
Suppose a population has 8,000 adults in Region A and 2,000 in Region B. A survey samples 400 from each region because the smaller region needs enough observations for reliable reporting.
The selection probability in Region A is 400 ÷ 8,000 = 0.05, giving a design weight of 20. In Region B it is 400 ÷ 2,000 = 0.20, giving a design weight of 5.
If we simply average the 800 respondents, Region B receives half the influence even though it contains only one fifth of the population. The design weights restore the intended population representation.
Statistics Canada’s weighting and estimation guidance makes this distinction explicit: each sampled unit has a design weight based on selection, and estimation weights may later incorporate nonresponse or calibration adjustments.
A weight is not a confidence score
A unit with weight 20 is not twenty times more trustworthy than a unit with weight 1. It represents more population units under the design.
Nor does a large weight mean the observation should be copied twenty times into a spreadsheet. Replication changes the apparent sample size and can produce incorrect standard errors if the software then treats the copies as independent observations.
The estimator should use the weights while preserving the original sampled units and the design information needed for variance estimation.
Why unequal selection can be deliberate
Oversampling a small subgroup can be efficient when subgroup estimates matter. Stratified sampling can also reduce variance when the strata are internally homogeneous. Cluster sampling can reduce fieldwork cost even though it often increases sampling variance.
The final weight is therefore not necessarily evidence that something went wrong. It may be the planned mathematical consequence of the design.
The mistake is to ignore the design at analysis time and treat every observed row as though it arrived with the same population meaning.
Nonresponse creates a second selection problem
Sampling selects people into the study. Response then selects a subset of those sampled units into the usable dataset. If response propensities differ systematically, the respondents can become less representative than the original sample.
Suppose each of two equally sized population groups contributes 100 sampled units. Ninety respond in Group 1 but only fifty in Group 2. If the original design weight is the same, analysing respondents without adjustment gives Group 1 almost twice as much influence as Group 2.
A simple response-class adjustment redistributes the weights of nonrespondents to respondents within each class. Group 1 receives an adjustment factor 100 ÷ 90 ≈ 1.11. Group 2 receives 100 ÷ 50 = 2.
This repair is justified only if respondents can represent nonrespondents sufficiently well within the adjustment classes for the target estimand. Statistics Canada’s educational explanation of weighting adjustments emphasizes the same idea: nonresponse weights redistribute the design weights of nonrespondents to respondents under an assumption about their similarity.
Response rate is not the same as nonresponse bias
A low response rate creates concern because there is more missing representation to explain. But bias depends on how response relates to the survey variables, not on the response rate alone.
A 60% response rate can produce little bias for one estimate if response is nearly unrelated to that outcome after adjustment. A 90% response rate can still bias a rare subgroup estimate if most of the missing 10% come from that subgroup.
That is why strong nonresponse analysis compares respondents and nonrespondents using frame or administrative information where available, examines response patterns across important groups, and reports which variables the adjustment model uses.
Calibration uses trusted population information
Suppose an external population source tells us the number of people by age group, region and sex. Calibration adjusts survey weights so weighted sample totals match selected known totals while keeping weights reasonably close to their starting values.
Statistics Canada’s guidance distinguishes design weights from calibration weights and explains why auxiliary information can improve estimation. Calibration is especially useful when the auxiliary variables are related to survey outcomes and are measured consistently in both the survey and population source.
The principle is simple: if the population total is known more reliably elsewhere, the survey should not contradict it merely because the sample composition drifted.
Poststratification is a special, visible form of calibration
Poststratification divides the sample into mutually exclusive cells with known population totals, then adjusts weights so each cell reproduces its population count.
If a sample contains 80 respondents in a group known to contain 1,600 people, a simple poststratification weight for that cell is 20. Another cell with 200 respondents representing 1,000 people receives weight 5.
The method is transparent but can become unstable when many cross-classified cells are sparse or empty. A four-way table across age, region, education and language can create hundreds of cells even when each individual variable has only a few categories.
Raking matches margins instead of every cross-classified cell
Raking, also called iterative proportional fitting in this context, repeatedly adjusts weights so survey margins match known population margins for several variables.
For example, the weighted sample may be adjusted to match the population age distribution, then region distribution, then education distribution, cycling until the margins converge sufficiently.
Raking can avoid the extreme sparsity of full poststratification, but it does not force every interaction among the variables to match the population. If the joint age-by-region structure matters strongly and the sample represents it poorly, matching separate age and region margins may be insufficient.
Calibration equations are promises about totals
In general form, calibration chooses adjusted weights wᵢ so that weighted auxiliary totals in the sample equal known population totals:
Σ wᵢxᵢ = X, where xᵢ is a vector of auxiliary values and X is the corresponding known population total vector.
Statistics Canada’s technical discussion of calibration weighting presents this structure and explains how calibrated weights can be close to original design weights while satisfying auxiliary-total constraints.
The equation is not magic. Its usefulness depends on the quality, relevance and comparability of X.
Bad auxiliary information can make a survey worse
If the external totals use different definitions, outdated classifications or incomplete coverage, forcing survey estimates to match them can import their defects.
Suppose a population register classifies residence by legal address while the survey asks where a person usually lives. Calibration to the register may appear precise while aligning two different concepts.
Before calibration, document the source, reference date, population definition, coding system and uncertainty of the auxiliary totals. The general infrastructure belongs to Metadata and Data Lineage.
Coverage error is not just nonresponse
A sampling frame can omit eligible units entirely or include ineligible or duplicated units. Nonresponse occurs after a unit is sampled. Undercoverage can prevent the unit from ever having a chance of selection.
Calibration may reduce coverage bias when undercovered units resemble covered units within informative auxiliary categories. But if a missing population has distinct outcomes and no adequate representation in the sample, reweighting cannot observe them into existence.
This is the same boundary encountered in How External Validity and Evidence Transfer Work: weighting can rearrange evidence across supported regions; it cannot validate extrapolation across an empty region by arithmetic alone.
Extreme weights reveal both correction and fragility
If a respondent receives a very large final weight, that person is carrying substantial representation. The large weight may be mathematically necessary because similar units were rarely selected or rarely responded.
It also increases variance. If one high-weight observation changes, the population estimate can move sharply.
Extreme weights are therefore a diagnostic. Ask why they exist: rare subgroup, low response propensity, aggressive calibration, frame problem or model instability?
Weight trimming trades variance for bias
Analysts sometimes cap or smooth very large weights to reduce variance. This can stabilise estimates but alters the estimator and can reintroduce bias.
There is no universal safe threshold. A cap that works for one outcome may distort another, especially when the high-weight cases belong to a substantively important rare group.
Any trimming rule should be declared, justified and tested through sensitivity analysis. Report the distribution of weights before and after adjustment and show whether key estimates change materially.
Effective sample size can be much smaller than the row count
When weights vary, the information content of a weighted sample can be lower than that of an equal-weight sample with the same number of rows.
A commonly used rough diagnostic based only on weight variability is Kish’s effective sample size:
n_eff = (Σwᵢ)² ÷ Σwᵢ².
If 100 observations have equal weights, n_eff is 100. If a few observations carry most of the weight, n_eff can be far smaller.
This formula is a diagnostic, not a complete variance estimator. Clustering, stratification, finite-population corrections, outcome relationships and the process used to estimate the weights can all matter.
Weighted means are ratios of weighted totals
A weighted mean is usually calculated as:
weighted mean = Σwᵢyᵢ ÷ Σwᵢ.
This seems obvious, but mistakes are common when analysts normalise weights inconsistently across subgroups, filter records after calibration or use software that interprets “weights” as frequency weights rather than survey probability weights.
Always identify what a software weight argument means. Probability, analytic, frequency and replication weights can trigger different calculations.
Variance must respect the sample design
Using the correct point-estimate weights but ordinary independent-observation standard errors can still produce misleading uncertainty.
Stratification can reduce variance. Clustering often increases it because observations within a cluster are correlated. Weight estimation can add uncertainty. Complex surveys therefore use design-based variance formulas, Taylor linearisation or replicate-weight methods such as jackknife, balanced repeated replication or bootstrap variants.
The exact method depends on the design. A public-use dataset may provide replicate weights precisely because the released microdata no longer contain all information needed to reconstruct the original sampling process safely.
Subpopulation analysis needs the full design information
Suppose the survey contains 2,000 people, but the question concerns 300 young adults. Deleting all other respondents before variance estimation can give the wrong design-based standard error because the original strata and clusters still shape the sampling variance.
Survey software often provides domain or subpopulation estimation so the target subgroup is analysed while retaining the full design structure.
This is another example of a recurring principle: the rows used in the numerator are not necessarily the whole information needed to calculate uncertainty.
A worked fictional calibration example
Imagine a survey whose respondents, after design and nonresponse adjustment, represent 10,000 adults. The current weighted totals are 6,000 in age group A and 4,000 in age group B. A trusted population source says the correct totals are 5,000 and 5,000.
A simple poststratification adjustment multiplies age-group-A weights by 5,000 ÷ 6,000 = 0.8333 and age-group-B weights by 5,000 ÷ 4,000 = 1.25.
Suppose the weighted outcome mean is 70 in Group A and 50 in Group B. Before calibration, the overall weighted mean is 0.6 × 70 + 0.4 × 50 = 62. After calibration it is 0.5 × 70 + 0.5 × 50 = 60.
The two-point shift does not mean the original responses changed. The population composition represented by those responses changed because the survey was brought into alignment with a trusted population margin.
Calibration can improve one variable more than another
Auxiliary variables reduce bias or variance most effectively when they are related to the survey outcomes and the response or coverage process.
Calibrating age and region can improve estimates of an outcome strongly related to age and region. It may do little for another outcome driven by an unmeasured characteristic.
That is why there is no single “representative weight” that guarantees every variable is unbiased. Survey quality is estimate-specific.
Weighting cannot repair a changed question
Suppose a survey asks only people who used a service about their satisfaction, then weights those respondents to the full population. No weighting scheme can turn service-user satisfaction into the opinions of people who never used the service.
The target variable was never observed for the nonusers under the same meaning. The problem is not merely unequal representation; it is a different estimand.
Define the population and outcome before constructing weights. Weighting is not permission to broaden the claim after the data are collected.
Propensity weighting is still model-based adjustment
Instead of broad response classes, analysts can model each sampled unit’s probability of response using observed auxiliary variables, then adjust weights inversely to estimated response propensity.
This allows more flexible adjustment but adds model dependence. Poorly specified propensity models can create extreme weights or fail to capture outcome-relevant differences between respondents and nonrespondents.
Cross-validation of predictive fit is not enough. A response model can predict response well using variables weakly related to the survey outcome while omitting a less predictive variable that matters strongly for bias. The model should be built for the estimation problem, not just classification accuracy.
One-step and two-step weighting answer different process questions
Some systems combine nonresponse and population calibration in one adjustment. Others first adjust respondents to represent the original sample, then calibrate the respondent-adjusted sample to population totals.
Statistics Canada’s research on one-step versus two-step calibration weighting shows that these approaches can behave differently under response and prediction models.
The practical lesson is not that one ordering always wins. It is that the sequence embodies assumptions about where representation was lost and how auxiliary information repairs it.
Weights need versioning
Weights can change when population benchmarks are revised, response models are improved, duplicate records are corrected or the sample frame is updated.
A published estimate should therefore preserve the weight version, benchmark source and reference date. Re-running an old dataset with new calibration margins can legitimately produce a different estimate.
That does not mean one version was fraudulent. It means the evidence object changed. The article How Timekeeping, Calendars and Time Standards Work addresses time identity generally; in survey work the same principle becomes a data-vintage question.
A weight-quality audit
- Define the target population and estimand.
- Document the sampling frame and selection probabilities.
- Compute or verify design weights.
- Describe eligibility changes and unknown eligibility.
- Examine unit and item nonresponse.
- Choose response-adjustment variables with substantive justification.
- Document auxiliary population totals and their source dates.
- Choose calibration or poststratification constraints.
- Inspect extreme and zero weights.
- Test sensitivity to trimming or model choices.
- Use variance estimation appropriate to the complex design.
- Preserve strata, clusters and replicate-weight information where needed.
- Report unweighted sample counts alongside weighted estimates.
- Keep the final claim inside the survey’s coverage and measurement boundaries.
What readers should ask when they see a weighted percentage
- What population does this percentage represent?
- How were sampled units selected?
- What are the largest and smallest weights?
- How was nonresponse handled?
- Which external totals were used for calibration?
- Do those totals use the same definitions and reference date?
- How was sampling variance calculated?
- Are important population groups absent or weakly represented?
- Would the conclusion change under a defensible alternative weighting model?
What this adds to the Library
A survey is not a bag of answers. It is a designed route from a population to a sample, from a sample to respondents, and from respondents back to population statements.
Weights record that route numerically. They preserve who was more likely to enter the data, who became underrepresented, and which trusted external totals were used to repair composition.
The most important boundary is simple: weighting can change how much influence observed evidence receives. It cannot change an unobserved population into observed evidence without assumptions that deserve to be named.
Sources and further reading
Statistics Canada materials were checked for this edition on 5 September 2026. The worked numerical examples are original and hypothetical.
- Statistics Canada — Weighting and Estimation
- Statistics Canada — Weighting
- Statistics Canada — Using calibration weighting to adjust for nonresponse and coverage errors
- Statistics Canada — One step or two? Calibration weighting from a complete list frame with nonresponse
Continue through eduKate: use Surveys and Sampling for design, Missing Data Analysis for incomplete variables, Small Area Estimation for borrowing strength across sparse domains, and Official Statistics for governance and public trust.