Research begins before a search box is opened. It begins when somebody notices that a question matters enough to deserve disciplined attention. A learner asks why two sources disagree. A scientist asks whether an observed effect is real. A historian asks what can be inferred from a surviving record. A parent asks whether a claim about learning is supported by evidence or merely repeated often enough to sound true. An engineer asks whether a failure was caused by design, materials, use, environment or some interaction among them.
The visible output of research may be a paper, graph, report, essay, dataset, policy recommendation, experiment, archive entry or classroom explanation. The invisible work underneath is more important: defining the question, choosing a method that can actually answer it, deciding what counts as evidence, checking where evidence came from, testing alternative explanations, estimating uncertainty, recording limitations and making sure the conclusion does not outrun what the evidence can support.
This is why research methods and source evaluation belong together. A strong source used with a weak method can mislead. A strong method applied to weak or inappropriate evidence can also mislead. Good research requires both a reliable route to evidence and disciplined judgement about what that evidence means.
The core research loop
A useful way to understand research is as a loop rather than a straight line:
QUESTION → DEFINE → SEARCH → SELECT → VERIFY → METHOD → OBSERVE OR MEASURE → ANALYSE → TEST ALTERNATIVES → ESTIMATE UNCERTAINTY → INTERPRET → REPORT → CHALLENGE → CORRECT → NEW QUESTION
Every stage can change the others. A search may reveal that the original question is too vague. A measurement problem may force a redesign. A conflicting source may expose an assumption. New evidence may make the initial hypothesis less plausible. Replication may show that an apparently strong result depends on a hidden condition. Research is therefore disciplined correction, not the ceremonial confirmation of what the researcher hoped to find.
1. Start with a question that can be answered
“What is the best education system?” sounds like a research question, but it hides several unresolved choices. Best for what outcome? Academic attainment, equity, creativity, wellbeing, social mobility, economic productivity or something else? At what age? Over what time period? Measured by which indicator? Compared across which populations? Under which constraints?
A researchable question reduces ambiguity without pretending the world is simple. It identifies the object of study, the comparison or relationship of interest, the relevant time and place, and the type of claim that could reasonably be supported.
- Descriptive questions ask what exists, how much, how often, where or when.
- Comparative questions ask how two or more cases differ or resemble one another.
- Explanatory questions ask why a pattern occurs or what mechanism produces it.
- Causal questions ask whether changing one factor changes another.
- Interpretive questions ask how people, institutions, texts or artefacts create and carry meaning.
- Historical questions ask what happened, how we know, and how the past developed through time.
- Design questions ask what intervention, system or structure is likely to achieve a specified outcome under constraints.
A major research error happens when a method designed for one kind of question is used to answer another. Correlation can describe association, but by itself it does not establish causation. A vivid interview can reveal experience and meaning, but it cannot automatically estimate prevalence in a population. A national average may show a broad pattern while hiding important differences among regions, schools, occupations or age groups.
2. Turn concepts into observable evidence
Many important concepts cannot be observed directly. Intelligence, trust, poverty, institutional quality, resilience, motivation, social cohesion and learning are not single physical objects that can be placed on a scale. Researchers therefore operationalise concepts: they decide which observations or measurements will stand as evidence for the thing being studied.
This translation from concept to measure is one of the most consequential parts of research. If the measure does not represent the concept well, a precise result may still answer the wrong question.
- Validity asks whether the method or instrument measures what it is intended to measure.
- Reliability asks whether a measurement is sufficiently consistent under appropriate repeated conditions.
- Sensitivity asks whether the method can detect a relevant change or difference.
- Specificity asks whether the method distinguishes the intended phenomenon from plausible alternatives.
- Resolution asks how much detail the measurement preserves.
- Calibration asks whether the measuring process is anchored to a trustworthy standard.
The United States Office of Research Integrity stresses that responsible data collection depends on appropriate, reliable methods and that bias or inappropriate procedures can compromise research data. The same underlying lesson applies far beyond laboratory research: a method is not good because it is sophisticated; it is good when it is fit for the question and transparent enough to be evaluated.
3. Know what kind of evidence you are looking at
Evidence takes many forms. Each form can be powerful when used for the job it can actually perform.
| Evidence form | Typical strengths | Typical limits |
|---|---|---|
| Direct observation | Close to the phenomenon; can reveal behaviour and context | Observer effects, limited coverage, interpretation risk |
| Experiment | Strong control over variables; can support causal inference under suitable design | May simplify reality; ethical or practical limits |
| Survey | Can cover large populations and standardised questions | Sampling, wording, recall and response biases |
| Interview | Depth, motivation, meaning, lived experience | Not automatically representative; interviewer and memory effects |
| Administrative data | Large scale, operational relevance, longitudinal potential | Collected for administrative purposes, not always for the research question |
| Historical record | Evidence from the period or institution being studied | Survival bias, authorship, motive, incomplete context |
| Material artefact | Physical evidence of manufacture, use, environment and custody | Provenance gaps, dating uncertainty, interpretation limits |
| Dataset | Allows systematic analysis and reuse | Definitions, missingness, collection method and metadata may constrain meaning |
| Model or simulation | Tests relationships, scenarios and mechanisms | Output depends on assumptions, structure and input quality |
| Secondary synthesis | Brings multiple studies or sources together | Depends on search strategy, inclusion criteria and underlying evidence |
The important question is not “Is this evidence?” but “Evidence for what claim, under what conditions, with what uncertainty?”
4. Source evaluation is claim evaluation
Students are often taught to judge a source by checking its author, date, publisher and domain. These checks are useful, but they are only the beginning. A source is not globally reliable or unreliable. It may be authoritative for one claim and weak for another.
A government statistical agency may be authoritative for an official population estimate but not for an independent moral judgement about public policy. A company may be the primary authority on the specifications of its own product but not a neutral authority on whether that product is the best choice. A newspaper may accurately report that an event occurred but may not provide enough methodological detail to evaluate a scientific claim mentioned in the story. A peer-reviewed article may contain serious limitations even though it passed review.
Source evaluation therefore works best when tied to the specific claim being supported.
CLAIM → WHAT EVIDENCE WOULD SUPPORT IT? → WHO COULD KNOW THIS? → HOW WOULD THEY KNOW? → WHAT METHOD DID THEY USE? → WHAT INCENTIVES OR CONSTRAINTS EXIST? → CAN THE EVIDENCE BE CHECKED? → DOES AN INDEPENDENT SOURCE AGREE? → WHAT REMAINS UNCERTAIN?
5. Primary and secondary sources do different jobs
A primary source is close to the event, experiment, institution, artefact or dataset being studied. A secondary source interprets, analyses or synthesises primary material. Neither category is automatically superior.
A primary document can reveal what an institution officially said, but not whether the statement was true. A laboratory paper can report an original experiment, but a later systematic review may provide a stronger estimate of the broader evidence. A diary can illuminate an individual’s experience while leaving the experience of an entire population unknown. A museum accession record can establish a custody event, while provenance research may be needed to reconstruct earlier ownership.
Good research often moves between levels: primary evidence for closeness and detail; secondary synthesis for context, comparison and cumulative judgement.
6. Authority matters, but method matters more
Institutional authority is useful because strong organisations often maintain expert review, documented procedures, quality control and accountability. Yet authority should not be used as a substitute for reading the evidence.
- Does the source show how the information was obtained?
- Are definitions visible?
- Is the relevant population or sample described?
- Are exclusions or missing data explained?
- Are uncertainty and limitations reported?
- Can another researcher, institution or reader inspect the route from evidence to conclusion?
- Has the source been corrected, updated or superseded?
- Is the publication speaking inside its area of competence?
The National Institutes of Health describes scientific rigor as the strict application of the scientific method to support unbiased and well-controlled design, methodology, analysis, interpretation and reporting. That emphasis on the full chain matters: a conclusion is only as trustworthy as the path that produced it.
7. Reliability is not the same as truth
A measurement can be reliable but invalid. A scale that is consistently five kilograms wrong may produce repeatable readings without producing accurate ones. A questionnaire can reliably measure test-taking confidence while being mistakenly interpreted as a direct measure of mathematical ability. A ranking system can consistently produce the same order while embedding a questionable definition of what counts as “best”.
Reliability asks whether a process behaves consistently. Validity asks whether the inference drawn from it is justified. Research needs both.
8. Bias enters before, during and after data collection
Bias is not simply a researcher having an opinion. In research, bias is a systematic process that pushes observations, measurements, selections, analyses or interpretations away from a fair representation of the phenomenon.
- Selection bias: the cases included differ systematically from the cases that matter.
- Sampling bias: the sample does not adequately represent the target population.
- Measurement bias: the instrument or procedure systematically distorts measurement.
- Recall bias: participants remember past events unevenly.
- Observer bias: expectations influence recording or interpretation.
- Confirmation bias: evidence supporting an expectation receives more attention than evidence against it.
- Publication bias: positive or striking results are more likely to be published than null findings.
- Survivorship bias: analysis focuses on what remains visible while missing what failed, disappeared or was never preserved.
- Historical source bias: surviving records disproportionately represent people or institutions able to create and preserve records.
- Algorithmic bias: data, labels, objectives or deployment conditions generate systematically unequal outcomes.
Bias cannot always be eliminated, but it should be anticipated, reduced where possible, measured when possible and declared when it remains.
9. Correlation, causation and mechanism
Two variables can move together because one causes the other, because the direction of influence runs the other way, because a third factor affects both, because of selection effects, because of measurement choices or because the apparent relationship arose by chance.
Causal research therefore asks for more than association. It looks for timing, plausible mechanisms, counterfactual reasoning, controls, natural or designed comparisons, robustness tests and alternative explanations. Different disciplines use different tools, but the underlying discipline is the same: do not upgrade an association into a causal claim without enough evidence.
10. Sampling determines what a study can speak about
A sample is a bridge from observed cases to a larger population. The bridge works only when the relationship between sample and target population is understood.
Large samples can reduce random error, but size does not repair systematic selection problems. Ten thousand voluntary online responses may be less representative of a population than a carefully designed probability sample of far fewer people. The researcher must therefore ask who had a chance to be included, who did not, who refused, who dropped out and whether those patterns matter to the result.
11. Transparency allows research to be challenged
Good research is not defined by never being wrong. It is defined partly by making the route to the result visible enough that errors can be discovered and corrected.
The National Academies’ work on reproducibility and replicability distinguishes computational reproducibility from the broader question of whether independent research can obtain consistent results. The larger lesson is that methods, data, code, assumptions and analytical choices should be described clearly enough for meaningful checking wherever ethical, legal and practical constraints allow.
- State the research question and hypotheses.
- Describe data collection and sampling.
- Define variables and measures.
- Report exclusions and missing data.
- Explain analytical choices.
- Distinguish planned analyses from exploratory ones where relevant.
- Report uncertainty rather than only point estimates.
- Preserve data, code, materials or provenance where lawful and appropriate.
- Record corrections and version changes.
12. Replication is a feature, not an insult
Research becomes stronger when important claims can survive independent checking. Replication may confirm a result, narrow the conditions under which it holds, expose a hidden dependency or reveal that the original finding was unstable. A non-replication does not automatically prove misconduct or incompetence; complex phenomena can vary across populations, settings, instruments and time.
What matters is whether the research community can learn from the difference. A mature research culture treats correction as part of knowledge production rather than as an embarrassment to be hidden.
13. Triangulation asks whether different routes converge
No single method is ideal for every dimension of a difficult question. Researchers often strengthen inference by triangulating across methods, datasets, institutions or forms of evidence.
For example, a study of urban transport reliability might combine operational records, passenger surveys, vehicle telemetry, maintenance logs, observation and policy documents. Agreement among independent routes can increase confidence. Disagreement can be even more useful because it identifies where definitions, perspectives or mechanisms need closer examination.
14. Uncertainty is information
Research claims are rarely all-or-nothing. Measurements have error. Samples vary. Models simplify. Sources are incomplete. Historical records contain gaps. Human testimony can be sincere and still imperfect. Predictions depend on assumptions.
A responsible conclusion therefore separates what is strongly supported, what is plausible, what is contested and what is unknown. Statistical uncertainty may be expressed through confidence or credible intervals, standard errors or sensitivity analyses. Qualitative research may express uncertainty through competing interpretations, source limitations, negative cases and explicit boundaries of inference. Historical work may distinguish documented fact from reconstruction or conjecture.
Uncertainty is not weakness. Hidden uncertainty is weakness.
15. A practical source-evaluation ladder
When a reader encounters an unfamiliar claim, the following ladder is useful:
- Identify the exact claim. Do not evaluate an entire article when only one sentence needs checking.
- Find the closest source. Look for the original dataset, study, document, law, standard, speech, archive record or institutional release.
- Check authority. Ask whether the source is in a position to know the claim.
- Check method. Ask how the evidence was produced.
- Check date and version. Determine whether the information is current for the claim.
- Check definitions. Many apparent disagreements are definition disagreements.
- Check independent support. Look for another reliable route to the same conclusion.
- Check incentives and omissions. Ask what the source gains, controls or leaves unexplained.
- Check scope. Determine whether the conclusion has been generalised beyond the population, time or setting studied.
- Record confidence. Decide what remains uncertain rather than forcing a false yes/no verdict.
16. Research in the humanities and history
Not all rigorous research looks like an experiment. Historians, literary scholars, archaeologists and researchers in the humanities often work with texts, artefacts, archives, images, material traces, language, institutions and cultural context.
Their methods may include source criticism, textual analysis, provenance reconstruction, dating, comparison, contextualisation and interpretation. A key discipline is to ask who produced the source, for whom, under what conditions, for what purpose, what it could observe, what it could not observe, why it survived and what other evidence supports or challenges it.
This is especially important because the archive is never a perfect recording of the past. What survives is shaped by power, institutions, storage, destruction, chance and later collection decisions.
17. Research in the social sciences
Social systems are difficult because human beings respond to institutions, incentives, expectations and one another. The act of measuring can sometimes change behaviour. Definitions vary across countries and time. Social categories can be historically contingent. Policies are rarely assigned randomly. Outcomes may have multiple causes operating at different levels.
Strong social research therefore pays close attention to case selection, measurement equivalence, confounding, institutional context, causal identification and the difference between individual-level and system-level conclusions. This is the bridge to eduKateSingapore’s World Knowledge Research Library, where cross-system claims should preserve definitions, context and ownership rather than flattening the world into a single ranking.
18. Research in the age of AI
AI changes the speed of research but does not remove the need for research judgement. A language model can help generate search terms, compare explanations, summarise a paper, classify evidence or suggest alternative hypotheses. It can also fabricate references, blur primary and secondary material, compress important uncertainty, reproduce bias or state an outdated claim fluently.
The correct research posture is therefore not “AI or sources”. It is AI routed through sources, methods and verification. The human or institutional research process remains responsible for identifying the claim, finding authoritative evidence, checking versions, preserving provenance and deciding what the evidence supports.
For eduKate, this is why retrieval readiness is not the same as publication readiness. A searchable page is useful only when the Library can know what it owns, what evidence supports it, when it was current and where the reader should go next.
19. Common research failure modes
- Starting with a conclusion and searching only for support.
- Using a proxy measure as though it were the concept itself.
- Confusing an official source with an independent source.
- Using a secondary article when the primary evidence is available.
- Ignoring publication date or version.
- Comparing statistics built from different definitions.
- Reporting averages without distributions.
- Treating correlation as causation.
- Using an impressive sample size to hide poor sampling.
- Ignoring missing data or attrition.
- Presenting model output without assumptions.
- Quoting evidence without reading the methods.
- Ignoring evidence that contradicts the preferred explanation.
- Hiding uncertainty to make the conclusion sound stronger.
- Copying the language of a source instead of reconstructing the reasoning.
20. A research workflow for learners
A student does not need a laboratory or research institute to work rigorously. A school-level research workflow can be simple and powerful:
ONE QUESTION → THREE POSSIBLE EXPLANATIONS → FIVE GOOD SOURCES → ONE PRIMARY SOURCE WHERE POSSIBLE → DEFINITIONS WRITTEN DOWN → EVIDENCE TABLE → CONTRADICTORY EVIDENCE → LIMITATIONS → BEST CURRENT CONCLUSION → WHAT WOULD CHANGE MY MIND?
The final question is especially important. If no possible evidence could change a conclusion, the activity has stopped being research and become defence of a belief.
21. A research workflow for institutions
Institutions need additional layers because knowledge must outlast individuals. The process should preserve source identity, version, method, ownership, permissions, corrections, unresolved disputes and update rules.
This is why eduKate separates acquisition, publishing, archives and data management. The Collections Development, Curation & Acquisitions Wing asks whether a real gap exists. Wintour House decides whether a manuscript deserves publication and whether its evidence, structure and edition are ready. The Research Collections Directory provides public research routes. How Data Management Works explains how evidence remains usable through capture, structure, validation, governance and preservation.
22. The difference between information and evidence
Information becomes evidence only in relation to a claim. A temperature reading is information. It becomes evidence when used to support a claim about fever, climate, equipment performance or chemical change, and the meaning depends on the measurement conditions. A photograph is information. It becomes evidence for a historical claim only after date, location, authorship, manipulation, subject and context are considered.
This distinction is central to critical thinking. The world contains more information than any person can inspect. Research is the discipline of deciding which information can legitimately bear the weight of which conclusions.
23. The strongest conclusion is often narrower
Weak research often tries to sound universal. Strong research frequently becomes more precise as it improves. Instead of “this teaching method works”, a better conclusion may be “under these conditions, for this group, using this outcome measure, the intervention produced an average improvement of this size, with these limitations”.
Narrower is not smaller when it is more accurate. Precision gives future research something solid to test.
24. Research ends by returning to the world
The purpose of disciplined research is not simply to accumulate citations. It is to improve what can be known, decided, taught, designed, preserved or questioned. A good research product therefore returns more than an answer. It returns a visible method, a source trail, an uncertainty boundary and a route for correction.
That is how a library becomes more than a collection of pages. Each article becomes a tested position in a larger knowledge system: one question answered as well as current evidence allows, connected to the evidence behind it, the systems around it and the next question it makes possible.
Sources and further reading
- National Institutes of Health — Enhancing Reproducibility through Rigor and Transparency
- National Academies — Reproducibility and Replicability in Science
- U.S. Office of Research Integrity — Data Collection
- Singapore Statement on Research Integrity
- NINDS — Rigorous Study Design and Transparent Reporting
eduKate route: Continue with How Scientific Research Works, Research Data Management and FAIR Principles, How Archives Work and the Research Collections Directory.
More articles in this collection
Research methods, evidence and inference
- How Case Study Research and Process Tracing Work | From Case Selection and Within-Case Evidence to Mechanisms, Rival Explanations and Causal Inference
- How Causal Inference Works | From Counterfactuals and DAGs to Experiments, Natural Experiments, Confounding and Treatment Effects
- How Censuses and Population Statistics Work | From Enumeration, Registers and Sampling to Demographic Evidence and Planning
- How Clustered and Multilevel Data Work | When Observations Belong to Groups, Places and Repeated Contexts
- How Construct Validity and Measurement Models Work | From Test Scores and Latent Concepts to Evidence, Fairness and Valid Use
- How Experimental Design Works | From Question and Randomisation to Controls, Blocking, Replication, Blinding and Causal Evidence
- How External Validity and Evidence Transfer Work | When Research Travels to a New Population
- How Forecasting and Prediction Work | From Baselines and Time Series to Probabilities, Calibration, Backtesting and Decision
- How Inter-Rater Reliability and Agreement Work | When Two People Judge the Same Evidence
- How Literature Search and Research Discovery Work | From Research Questions, Keywords and Controlled Vocabularies to Database Search, Citation Chaining, Alerts and Evidence Retrieval
- How Measurement Error and Misclassification Work | When the Recorded Value Is Not the Thing Itself
- How Missing Data Analysis Works | Missingness, Imputation and Honest Conclusions
- How Multiple Testing and Sequential Analysis Work | Many Questions, Repeated Looks and Honest Error Control
- How Observational Studies Work | From Cohorts and Cross-Sections to Bias, Confounding, Causal Limits and Real-World Evidence
- How Official Statistics Work | From Statistical Law and Professional Independence to Methods, Revisions, Metadata and Public Trust
- How Sensitivity Analysis and Robustness Checks Work | Which Assumptions Change the Answer?
- How Statistical Inference and Uncertainty Work | From Samples and Estimates to Confidence Intervals, Hypothesis Tests, P-Values and Bayesian Reasoning
- How Statistical Power and Sample Size Planning Work | From Research Question to Detectable Effect, Precision, Attrition and Honest Design
- How Surveys and Sampling Work | From Target Population and Questionnaire Design to Weighting, Nonresponse, Uncertainty and Inference
- How Time-to-Event and Survival Analysis Work | Censoring, Risk Sets, Kaplan–Meier Curves and Hazards
