A photograph. A laboratory reading. A witness statement. A government record. A spreadsheet. A broken component. A satellite image. A student’s working. A timestamp. A silence in an archive.
Which of these is evidence?
The answer is not “all of them” and not “none of them”. The answer is: evidence for what?
A photograph showing rain on a street can be evidence that rain was falling at a particular place and time if the image’s provenance is reliable. It is weak evidence for the annual rainfall of the city. The same image may be strong evidence that one road was flooded, weak evidence that every road was flooded, irrelevant evidence for tomorrow’s weather, and misleading evidence if it was taken three years earlier.
Evidence is therefore not merely a pile of information. It is a relationship between information and a claim.
The quality of reasoning depends not only on having evidence, but on asking what the evidence actually bears on, how it was produced, how independent it is, what uncertainty it carries, and where its support stops.
The answer in one paragraph
For practical reasoning, evidence is an observation, record, measurement, testimony, artifact, dataset or other information that bears on whether a particular claim should be supported, contradicted or left unresolved. Evidence is relational: the same item can be relevant to one claim and irrelevant to another. Good evaluation asks whether the evidence is relevant, sufficiently reliable for the task, independent of other evidence where independence matters, appropriately scoped in time and population, and strong enough—individually or in combination—to justify the conclusion being drawn. Evidence can increase confidence without proving certainty, and contradictory or missing evidence should be preserved rather than quietly averaged away.
This article follows What Is a Claim? | From Sentence to Verifiable Unit. That article asks what exactly is being asserted. The present article asks what can legitimately bear on that assertion. The next layer is the already-published What Is a Fact? | Claims, Evidence, Verification, Context and Change, which asks what conclusion is justified after the evidence has been examined.
The concept of evidence has a long philosophical history and no single definition settles every domain. The Stanford Encyclopedia of Philosophy entry on evidence surveys major debates and emphasises the close relationship between evidence and justified belief. This article adopts a deliberately practical, cross-domain grammar. It does not claim that science, law, history, auditing, medicine and everyday reasoning use identical standards.
1. Evidence is relational
A fingerprint is not simply “evidence” in the abstract. It is evidence relative to a proposition.
A fingerprint on a glass may bear on the claim that a person handled the glass. It does not by itself establish when the glass was handled, why it was handled, whether the person entered a room unlawfully, or whether the person committed another act nearby.
This gives us the first rule:
Never ask only “Is this evidence?” Ask “What claim is this evidence for or against?”
That small change prevents a large class of reasoning errors.
2. Data is not automatically evidence
Data are recorded values, observations, symbols or measurements. They become evidence when they are connected to a question or claim through an appropriate method.
A list of 10,000 temperatures is data. If the question is whether a particular room exceeded 30°C during an examination, only the observations tied to that room, time interval, sensor and measurement method may be relevant evidence.
More data can therefore produce less clarity if the added records do not belong to the evidential problem.
Evidence begins when data acquire a relationship to a claim.
3. A source is not the same as evidence
A book, website, interview, database or document is a source. The information within it may supply evidence for particular claims.
One source may support one sentence and contradict another. A source may be authoritative for its own rules yet weak for evaluating the effects of those rules. A company can be the best source for what it announced and a poor independent source for whether its product outperforms competitors.
“Good source” is incomplete language. The better question is: good source for which claim?
4. A citation is not the same as evidence
A citation points from one information object to another. It can help a reader locate supporting material. It does not guarantee that the cited source supports the sentence beside it.
A citation can fail through entity mismatch, time mismatch, population mismatch, quotation drift, selective reading, a stronger paraphrase than the source permits, or simple irrelevance.
The correct test is claim-to-evidence alignment: does this source, under this interpretation, actually bear on this exact claim?
5. Observation is not yet conclusion
Suppose a student observes droplets on the outside of a cold bottle.
“Droplets are visible on the outside surface” is an observation. “The water came through the bottle wall” is an explanation. “The droplets condensed from water vapour in the surrounding air” is another explanation.
The observation is evidence that any adequate explanation must account for the droplets. It does not choose the correct explanation by itself.
Evidence and inference are therefore connected but distinct.
6. Evidence can support, contradict or discriminate
Evidence does more than “support”. It can play several roles.
- Support: makes a claim more reasonable.
- Contradict: bears against a claim.
- Constrain: rules out part of the possibility space.
- Discriminate: helps choose between competing explanations.
- Calibrate: reveals how accurate a method or model tends to be.
- Contextualise: changes how another observation should be interpreted.
A powerful piece of evidence is often powerful because rival explanations predict different observations. Evidence that merely fits every live explanation may add little.
7. Relevance comes before prestige
A highly prestigious source can be irrelevant to the claim being tested.
Suppose an excellent national report describes average school attendance across a country. Your claim concerns whether one named lesson occurred on Tuesday at 3 pm. The national report may be impeccable and still contribute almost nothing to that question.
Relevance asks about the relationship between evidence and assertion.
In auditing, the US Public Company Accounting Oversight Board’s AS 1105 on Audit Evidence makes this relationship explicit within its own domain: appropriateness concerns the quality of audit evidence and includes relevance and reliability. That is an auditing standard, not a universal law of evidence, but the distinction is broadly instructive.
8. Reliability asks whether the evidence deserves the weight placed on it
Evidence can be relevant but unreliable.
A witness may have been too far away. A sensor may be uncalibrated. A database may contain stale records. A photograph may be authentic but misdated. A document may have been altered. A survey question may lead respondents. A model may be poorly validated for the population where it is being used.
Reliability is not a permanent property of a source brand. It depends on source, method, controls, context and the specific claim.
9. Reliability is not binary
It is tempting to divide evidence into reliable and unreliable piles. Real reasoning is usually more graded.
A thermometer with known uncertainty can still provide useful evidence. A memory can be imperfect yet informative. An old map may be unreliable for current roads and valuable for historical land use.
Ask how much weight the evidence deserves for the exact inference under consideration.
10. Sufficiency is not just quantity
Ten weak copies of one source are not necessarily stronger than one well-designed independent measurement.
Within auditing, AS 1105 distinguishes sufficiency, associated with quantity, from appropriateness, associated with quality, and explicitly notes that obtaining more of the same poor-quality evidence does not necessarily compensate for poor quality.
That is a useful warning beyond auditing: evidence quantity and evidence quality are different dimensions.
11. Independent evidence is different from repeated publication
Five websites can be one evidence lineage.
If four articles repeat a fifth article, which repeats one press release, the number of webpages has increased. The number of independent observations may still be one.
Independent corroboration occurs when evidence arrives through sufficiently separate pathways that the same error is less likely to explain all of it.
This is why provenance matters.
12. Provenance tells us how evidence came to exist
Provenance records the history of an information object: what generated it, what it was derived from, which agents were involved, which transformations occurred and which version is being inspected.
The W3C’s PROV-O Recommendation supplies a general vocabulary for describing entities, activities, agents, derivations, quotations, revisions, primary sources and other provenance relationships.
Provenance can reveal that two charts come from the same dataset, that a quotation came through an intermediary, or that a document is a revision of an earlier version.
But provenance is not truth. A beautifully documented error remains an error. Provenance helps us understand the lineage through which evidence should be interpreted.
13. Corroboration means compatible evidence from another route
Suppose a fire alarm log shows an alarm at 10:03. A security camera independently shows people leaving at 10:04. A witness remembers hearing the alarm shortly before evacuation.
The records are not identical. They arise through different processes and can jointly support a timeline.
Corroboration is strongest when the evidence pathways have different failure modes. Two sensors of the same type sharing the same broken calibration may agree and still be wrong together.
14. Triangulation is not majority voting
Triangulation combines evidence produced through different methods, perspectives or data sources to see whether a conclusion survives variation in how the question is approached.
It does not mean that three sources automatically defeat two.
A high-quality measurement may outweigh many informal impressions. A newly authenticated document may overturn a widely repeated historical assumption. An experiment may conflict with observational data for reasons that require deeper analysis.
The goal is not to count sources. It is to understand why they agree or disagree.
15. Direct and indirect evidence depend on the claim
The labels direct and indirect vary by domain, so they should not be treated as one universal hierarchy.
In everyday reasoning, a camera recording of a door opening may be relatively direct evidence that the door opened. A wet floor and an open umbrella nearby may be indirect evidence that somebody recently came in from rain.
Indirect evidence is not automatically weak. A carefully assembled network of indirect observations can support a conclusion powerfully. Direct-looking evidence can be deceptive if identity, time or provenance is wrong.
16. Positive evidence and negative evidence
Positive evidence records something observed: a detected signal, a document entry, a measured value, an event occurrence.
Negative evidence concerns a meaningful non-detection or absence under conditions where detection was expected.
“The search returned no result” becomes evidence of absence only when the search boundary, sensitivity and completeness are strong enough for that inference.
This connects directly to What Does a Blank Actually Mean?: no observation is not automatically an observation of none.
17. Absence of evidence is not always evidence of absence
If you search one drawer for a key and do not find it, you have evidence that the key is not in the part of the drawer effectively searched. You do not yet have strong evidence that the key is nowhere in the building.
If a complete inventory system is known to record every issued key and contains no record of one ever being issued, the evidential situation is different.
Negative evidence is therefore inseparable from coverage and detection capability.
18. Measurement produces evidence with uncertainty
A measurement is not made stronger by pretending uncertainty does not exist.
NIST’s Technical Note 1297 provides guidance for evaluating and expressing measurement uncertainty. NIST’s policy materials emphasise that measurement results should be accompanied by appropriate uncertainty information.
This is a profound evidential lesson. “100.000” can look more certain than “about 100 ± uncertainty”, while the latter may be the more honest scientific statement.
Evidence does not become weaker when its uncertainty is reported accurately. Its limits become visible.
19. Error models belong beside measurements
Every measurement process has ways it can go wrong.
- instrument drift;
- rounding;
- sampling error;
- operator variation;
- environmental interference;
- misclassification;
- data-entry error;
- missing observations;
- calibration error.
An evidence record becomes more useful when its likely error mechanisms are known. Reliability is not “I trust the instrument”. It is a reasoned understanding of how the instrument and procedure behave.
20. Evidence strength is claim-specific
A school attendance register can be strong evidence that an attendance mark was recorded. It may be moderate evidence that a student was physically present for the whole lesson. It is weak evidence that the student was cognitively engaged.
The evidence item has not changed. The target claim has.
This is why evidence quality cannot be assigned once and reused blindly across questions.
21. Evidence for identity is different from evidence for state
Before evidence about an entity can be combined, identity must be resolved.
Two records saying “A. Tan” received different scores may describe the same student, different students, or one student under two identifiers. If identity is unresolved, combining the scores can manufacture a false history.
Evidence can answer “What happened to this entity?” only after the “Which entity?” problem is sufficiently controlled.
Use When Are Two Things the Same Thing? for that layer.
22. Evidence for change needs two comparable states
To show that something changed, we need observations that describe the same relevant object or population at different times under sufficiently comparable definitions.
A score of 70 this year and 60 last year is not automatically evidence of improved mastery if the tests, scales, cohorts or conditions differ materially.
Change evidence requires identity, time, measurement comparability and uncertainty to line up.
23. Evidence for causation is different from evidence for sequence
“B happened after A” is evidence of temporal order. It is not by itself sufficient evidence that A caused B.
Causal evidence may require experiments, natural experiments, counterfactual reasoning, process evidence, longitudinal structure, mechanistic evidence or other designs appropriate to the domain.
The crucial point is not that one design is always superior. It is that causal claims demand evidence that addresses alternative explanations.
24. Statistical evidence is not just a p-value
Statistical evidence can involve effect sizes, uncertainty intervals, model fit, likelihoods, posterior distributions, predictive performance, sensitivity analyses and more.
No single number carries the whole evidential burden.
The design that generated the data, the population, missingness, model assumptions, multiplicity, measurement quality and practical importance all affect interpretation.
A tiny p-value attached to a badly posed question does not rescue the question.
25. Model-based evidence carries model dependence
Models transform observations under assumptions.
NIST’s current work on evidential statistics in forensic science highlights how evidential interpretations such as likelihood ratios depend on modelling choices, data processing and expert judgement.
The lesson is broader than forensics: a model output is evidence mediated through a model. Its weight depends partly on whether the model is appropriate, calibrated and robust for the case at hand.
26. Scientific evidence is a system, not a single experiment
One experiment can be important. Mature scientific confidence usually emerges from a wider structure: theory, measurement, replication, alternative explanations, independent laboratories, converging methods, correction and cumulative synthesis.
The relevant question is rarely “Is there a study?” It is “What does the body of evidence justify?”
This is why individual studies and evidence synthesis should not be conflated.
27. Evidence synthesis asks a different question
A systematic review does not merely collect papers. It defines a question, identifies eligible studies, assesses their methods and combines or compares findings under explicit rules.
Cochrane’s handbooks make these procedures explicit for health research. For example, review scope, inclusion criteria, study limitations and synthesis methods are defined rather than left implicit.
For a fuller treatment, see How Systematic Reviews and Evidence Synthesis Work.
28. Certainty in a body of evidence is not the same as number of studies
Cochrane uses the GRADE approach for bodies of evidence concerning intervention outcomes. Its current handbook describes considerations including risk of bias, inconsistency, indirectness, imprecision and publication bias.
That framework is domain-specific and should not be pasted onto every knowledge problem. Its deeper lesson is transferable: a body of evidence can be large yet uncertain because the studies share limitations, disagree, address the wrong population or remain imprecise.
Quantity does not erase structure.
29. Qualitative evidence answers questions numbers may not
Interviews, observations, field notes and other qualitative materials can provide evidence about experiences, meanings, mechanisms, implementation and context.
They should not be judged as failed numerical studies. They answer different questions through different methods.
Cochrane’s qualitative evidence synthesis guidance explicitly recognises the value of qualitative evidence for understanding intervention complexity, context, implementation and stakeholder experiences.
The evidential standard must fit the question.
30. Historical evidence lives inside incomplete archives
Historians rarely receive a complete recording of the past. They work with surviving letters, official records, newspapers, objects, photographs, oral histories, maps and later accounts.
Survival is selective. Powerful institutions may leave abundant records while marginalised communities leave fewer formal documents. Archives can preserve the bureaucracy of an event better than the lived experience of it.
Historical evidence therefore demands provenance, context, authorship, purpose, silence and comparison among sources.
For the specialist branch, see How Archival Evidence Works.
31. Legal evidence is governed by jurisdiction and procedure
Law uses formal rules about admissibility, relevance, testimony, authentication, burdens and standards of proof. Those rules vary by jurisdiction and proceeding.
Therefore the everyday phrase “legal evidence” should not be converted into a universal evidence hierarchy.
A piece of information can be factually informative yet inadmissible under a particular legal rule, or admissible without being conclusive. Legal admissibility and epistemic strength overlap but are not identical concepts.
32. Audit evidence connects information to an opinion
Auditing gives us a useful worked example of evidence as disciplined support for a conclusion.
PCAOB AS 1105 defines audit evidence as information used by the auditor in arriving at conclusions underlying the audit opinion, including information that corroborates and information that contradicts management assertions.
The inclusion of contradictory information is important. Mature evidence systems do not collect only confirming material.
33. Medical evidence has its own decision context
Clinical evidence interacts with diagnosis, prognosis, treatment effects, harms, patient values and individual circumstances.
A population-average treatment effect does not mechanically determine what should happen to one patient. The evidence must be interpreted through the clinical question and patient context.
This is why the dedicated What Is Medical Evidence? page remains the owner of that domain. The present article supplies the general grammar only.
34. Educational evidence must separate learning from proxies
A student’s test score is evidence of performance on a particular assessment under particular conditions. It is not the entire student, and it is not automatically a pure measure of underlying mastery.
Homework completion can be evidence of work submitted. Attendance can be evidence of recorded presence. A correct answer can be evidence of successful performance. None of these alone reveals every mechanism behind the result.
Educational reasoning becomes more accurate when proxies are named as proxies.
35. Test evidence should be matched to the construct
If a question is intended to measure algebraic reasoning but depends heavily on unfamiliar vocabulary, performance may reflect both mathematics and language access.
Evidence from assessment therefore raises a construct question: what does the result actually tell us about the ability we care about?
This is why validity is not merely a property printed on a test. It concerns whether evidence supports the intended interpretation and use of scores.
36. Images are evidence only after time, place and provenance are controlled
An image can show visible content vividly while hiding context.
To use an image as evidence, ask:
- When was it captured?
- Where?
- By whom or what system?
- Is it original or transformed?
- What is outside the frame?
- Does the caption identify the scene correctly?
- Does the image show the event claimed, or merely something similar?
Visual vividness should not be confused with evidential completeness.
37. Maps are evidence through a model of space
A map is not the territory. It is a representation created through coordinate systems, projections, classification choices, scale and data collection.
A road map may be excellent evidence for route structure and poor evidence for terrain elevation. A satellite-derived land-cover map may depend on classification algorithms and acquisition dates.
Evidence interpretation improves when the representation method remains visible.
38. Testimony is evidence with human conditions
People observe from positions. They remember selectively. They may be sincere and mistaken, informed and biased, accurate on one detail and uncertain on another.
Evaluating testimony can involve opportunity to observe, contemporaneity, consistency, incentives, expertise, independence and corroboration.
Credibility should not be reduced to whether a witness “sounds confident”. Confidence and accuracy are not interchangeable.
39. Documents are evidence of more than their contents
A document may provide evidence that:
- a statement was written;
- a policy existed in a particular version;
- a transaction was recorded;
- an institution intended something;
- a person received or sent a communication;
- a particular language or classification was used.
It may not establish that every statement in the document is true.
Authenticity, content truth and effect are separate questions.
40. Databases are evidence through data-generating processes
A database feels objective because it is structured. Structure does not remove upstream choices.
Ask:
- Who enters records?
- What counts as an event?
- Which fields are mandatory?
- What happens when data are missing?
- Are records corrected?
- Are deleted records recoverable?
- What timestamps mean creation, occurrence, update or ingestion?
- Which entities share identifiers?
A database is an evidence system with rules, not a transparent window onto reality.
41. Sensor evidence includes the sensor’s operating envelope
A sensor reading is meaningful only within a known measurement context.
Range, calibration, resolution, sampling frequency, placement, environmental sensitivity and maintenance can matter.
A sensor can accurately report its own internal reading while the reading poorly represents the physical quantity we intended to measure.
42. Model outputs are evidence about the model first
If a model predicts a 70% probability, one immediate fact is that the model produced that output under specified inputs and version.
Whether the 70% is well calibrated for the target world is a separate question.
Validation data, out-of-sample performance, calibration, subgroup performance and domain shift determine how much evidential weight the output deserves.
43. Expert judgement is evidence, not omniscience
Expertise can improve interpretation because experts know mechanisms, measurement limitations, domain conventions and relevant alternative explanations.
Expert judgement still has provenance and uncertainty. Experts can disagree. Incentives matter. Expertise can be strong in one subfield and weak in another.
The right question is not “expert or evidence?” Expert judgement is itself one form of evidence or evidential interpretation whose weight depends on the task.
44. Consensus is evidence about a community’s judgement
Consensus among qualified experts can be highly informative, especially when it emerges from independent engagement with a mature evidence base.
Consensus is not identical to truth. It can change as evidence changes.
A careful statement distinguishes “the current expert consensus is X” from “X is true because experts agree”. The former reports an epistemic state. The latter skips the underlying evidence.
45. Primary and secondary are not universal quality grades
A primary source is often close to an event or original research. A secondary source may synthesize, interpret or review primary material.
Closer is not automatically better.
An eyewitness may be close and mistaken. A systematic review may be secondary and far more informative for a treatment-effect question. An official primary record may establish what an institution recorded while an independent secondary investigation establishes that the record was incomplete.
Source type is one dimension, not a complete ranking.
46. There is no universal evidence pyramid
Different questions require different evidence structures.
A randomised trial may be powerful for estimating one kind of causal treatment effect. It may be useless for determining who authored a nineteenth-century letter. A satellite image can be excellent for land-cover change and irrelevant to a person’s private intention. A legal document can establish a rule and say little about how the rule affects outcomes.
The better architecture is question → claim → appropriate evidence, not one pyramid for every domain.
47. Contradictory evidence should stay visible
When two credible sources conflict, the system has learned something important: the evidence state is not yet coherent.
Do not immediately average the values or choose the source with the nicer website.
Investigate:
- same entity?
- same time?
- same definition?
- same measurement method?
- same population?
- one source derived from the other?
- one superseded?
- known error?
- genuine unresolved disagreement?
Contradiction is not noise to remove. It is a diagnostic signal.
48. Supporting evidence and contradicting evidence belong in the same frame
Confirmation bias grows when a search is designed only to find support.
A high-standard evidence review asks both:
- What would support this claim?
- What would we expect to see if this claim were wrong?
This creates an active search for falsifying or discriminating information rather than a passive collection of friendly sources.
49. Missing evidence can change the conclusion
If a dataset contains results for only the highest-performing students, the observed average cannot automatically describe the whole class.
If an archive contains only one side of a correspondence, interpretation changes. If sensor data disappear exactly during peak loads, the missing interval may be more important than the periods that remain.
Evidence evaluation must therefore include the pattern of what is missing, not merely the quality of what is present.
50. Freshness matters when the world can change
An official page from five years ago can be authentic, authoritative for its historical state and wrong for today’s question.
Claims about office holders, prices, opening hours, software behaviour, regulations, policies, availability and operating status need evidence from the relevant time.
Freshness is therefore part of relevance, not a decorative timestamp.
51. Evidence transfer across populations can fail
Evidence collected in one population, setting or period may not transfer unchanged to another.
A teaching intervention tested in one age group may behave differently in another. A medical study may underrepresent a subgroup. A transport model calibrated in one city may fail in a city with different travel patterns.
This is the problem of indirectness or external validity. The evidence can be high quality for its original question and limited for the new one.
52. Evidence can be precise and still answer the wrong question
Suppose a survey precisely estimates how satisfied current users are. The decision concerns why former users left.
The estimate may be statistically excellent and decision-relevant evidence may still be missing.
This is why relevance comes before technical sophistication.
53. Evidence standards should rise with consequence
Not every claim requires the same evidential burden.
Choosing a café can tolerate informal reviews and reversible error. Diagnosing disease, accusing a person of wrongdoing, changing public policy or making a high-value investment requires stronger safeguards.
Higher consequence usually justifies greater source independence, stronger methods, explicit uncertainty and more careful review.
Evidence thresholds are therefore connected to decision cost, not only abstract truth.
54. More evidence is not always worth collecting
Evidence gathering has costs: time, money, privacy, delay, opportunity and sometimes physical intervention.
If a decision would be the same across every plausible result of another test, more evidence may have little value.
If one discriminating observation could reverse a high-stakes decision, collecting it may be extremely valuable.
This connects evidence evaluation to How Value of Information Works: ask which unanswered question could actually change the decision.
55. Bayesian updating makes the relational idea explicit
Bayesian reasoning formalises one way evidence can change belief. A prior state of belief is updated according to how expected the evidence would be under competing possibilities.
The important intuition is not “use Bayes for everything”. It is that evidence has strength because of how differently it is predicted by alternatives.
Evidence that is equally likely under two explanations does little to choose between them. Evidence much more expected under one explanation can be strongly discriminating.
For the formal branch, see How Bayesian Inference and Updating Work.
56. Likelihood is not posterior probability
A common reasoning error is to confuse the probability of evidence given a hypothesis with the probability of the hypothesis given the evidence.
For example, a test result may be common when a condition is present while the condition itself remains uncommon in the tested population. The evidential interpretation requires both the test characteristics and relevant prior information.
The same caution appears in forensic evidential statistics: statistical evidence does not eliminate modelling assumptions or the need to communicate limitations.
57. Evidence can update without settling
Evidence does not need to produce certainty to be useful.
A new observation may move a claim from plausible to probable, or from probable to disputed. A better measurement may narrow an interval without fixing an exact value. An additional witness may strengthen one timeline while leaving motive unresolved.
Reasoning systems should therefore support graded states rather than only true/false.
58. Evidence does not justify stronger wording than it supports
Evidence says “associated”. A summary writes “caused”.
Evidence says “may”. A headline writes “will”.
Evidence says “in this sample”. A paragraph writes “people”.
Evidence says “no statistically significant difference detected”. A summary writes “the treatments are identical”.
Evidence discipline is partly a writing discipline: the verb, quantifier, population and uncertainty should remain faithful to what the evidence earns.
59. Evidence should be attached to the smallest meaningful claim
A paragraph may contain ten factual claims. One citation at the end can create false confidence because the reader cannot tell what the source actually supports.
The preceding article on claim extraction solves the first half of this problem by decomposing language into checkable units.
Evidence mapping solves the second half: each material claim is linked to the evidence that supports, contradicts or leaves it unresolved.
60. A practical claim–evidence matrix
| Claim | Evidence | Relation | Main limit | Status |
|---|---|---|---|---|
| The lesson was scheduled for 3 pm | Official timetable version 4 | Directly supports scheduled time | Does not establish actual start | Supported |
| The lesson began at 3 pm | Timetable | Indirect/weak | Plan is not outcome | Unresolved |
| Students were in the room at 3:04 | Authenticated photo at 3:04 | Supports visible presence | Frame incomplete | Supported within frame |
| The teacher entered at 3:06 | Access-control record | Supports recorded entry event | System may not record all entry routes | Conditionally supported |
| The delay caused lower learning | Same records | Insufficient for causation | No causal design | Unsupported by current evidence |
The matrix prevents one piece of evidence from silently doing five jobs.
61. Worked example: did a student understand the concept?
A student gets one question correct.
What does this prove?
It is evidence that the student produced a correct response under those conditions. It may support understanding. But guessing, memorisation, cue recognition or a familiar procedure may also explain the answer.
Add a second problem with changed numbers. Then a transfer problem in a new context. Ask the student to explain the reasoning. Ask for an error check. Evidence now arrives through different task demands.
Understanding becomes better supported because competing explanations have been tested.
62. Worked example: did the treatment work?
A patient improves after treatment.
This is evidence of improvement following treatment. It is not automatically sufficient evidence that the treatment caused the improvement. The condition might fluctuate naturally, another intervention may have occurred, measurement may have changed or regression to the mean may matter.
Clinical causal evidence therefore relies on study designs that address relevant alternative explanations and on bodies of evidence rather than one before-and-after story.
63. Worked example: did the policy reduce congestion?
Traffic volume falls 8% after a policy starts.
Evidence for what?
- Evidence that measured traffic volume changed after implementation.
- Not yet sufficient evidence that the policy caused the change.
- Potentially confounded by fuel prices, school holidays, roadworks, weather, remote work or other interventions.
- Dependent on where, when and how traffic was measured.
A causal design must separate the policy effect from other changes occurring at the same time.
64. Worked example: does the archive prove nobody objected?
An archive contains no letters opposing a proposal.
Possible interpretations include:
- nobody objected;
- people objected verbally;
- letters were never preserved;
- the relevant collection is incomplete;
- opposition appears under different terminology;
- the search strategy missed it.
Archival silence becomes stronger negative evidence only when the record-generating and preservation system makes the missing item genuinely expected.
65. Worked example: five news articles report the same number
Five articles say a project cost $800 million.
A provenance check shows all five quote one press release. The apparent five-source corroboration collapses to one upstream claim.
Then an audited filing gives $620 million for construction cost, while the press release’s $800 million includes land, financing and programme costs.
The contradiction was partly definitional. The evidence becomes coherent only after the cost boundary is typed correctly.
66. Worked example: an AI answer with three citations
An AI writes four factual sentences and provides three citations.
Verification should not ask merely whether the citations are reputable. It should extract the claims and check each relationship:
- Does citation A support claim 1?
- Does citation B support all of claim 2 or only the first clause?
- Is claim 3 an inference not stated by any source?
- Is claim 4 current, while citation C describes an older state?
An answer can have real citations and still contain unsupported synthesis.
67. AI-generated text is not evidence merely because it sounds informed
An AI response is an information object generated by a model. It can be useful for hypotheses, explanations, retrieval plans, transformations and synthesis.
Its prose is not automatically evidence for the factual claims it contains.
For factual use, the answer should route back to inspectable evidence whose relevance, provenance and limits can be evaluated.
This article describes a public reasoning pattern. It does not claim that any particular deployed AI system implements the pattern or achieves a measured reliability improvement.
68. Retrieval is not verification
Search retrieves potentially relevant material. Verification evaluates whether that material actually supports the claim.
Top-ranked results can repeat the same upstream error. Search snippets can omit qualifiers. A retrieved page can be about a namesake. A page can be current while the quoted statistic is old.
Retrieval is the beginning of the evidential path, not the verdict.
69. Search failure is evidence about the search
“I searched and found nothing” should initially be represented as:
No qualifying result was found under this query, source set, time and search procedure.
Only after search coverage is understood should that result be promoted into a stronger negative claim.
70. Evidence has a valid time
Some evidence remains relevant for centuries. Other evidence decays quickly.
A chemical constant may be stable. A restaurant opening hour may change tomorrow. A software interface may change next month. A regulation may be amended. A person’s role may end.
An evidence system should therefore distinguish:
- event time;
- observation time;
- record creation time;
- publication time;
- retrieval time;
- validity period.
Fresh-looking pages can contain stale facts. Old documents can be the best evidence for old states.
71. Evidence should preserve the route from world to record
A useful provenance chain often looks like:
world state → observation → measurement or record → transformation → publication → retrieval → interpretation → claim evaluation
Every arrow is a possible place for distortion.
The purpose of provenance is not bureaucratic completeness. It is the ability to locate where uncertainty, transformation or error entered the chain.
72. An evidence ledger should record contradiction, not just support
A practical evidence record can include:
- claim ID — which proposition is being evaluated;
- evidence item — observation, source, measurement or artifact;
- relation — supports, contradicts, constrains, contextualises;
- source — where the item came from;
- provenance — lineage and transformations;
- valid time — when the evidence applies;
- scope — population, location, jurisdiction or version;
- method — how the evidence was generated;
- uncertainty — quantitative or qualitative limitations;
- independence family — whether other items share the same upstream source;
- status — active, superseded, disputed, retracted or otherwise qualified.
This is a public information-design pattern, not a description of any proprietary eduKateAI implementation.
73. Evidence graphs preserve dependency
A flat list says that five items exist. An evidence graph can show that items 2, 3 and 4 derive from item 1, while item 5 is independent.
It can also show:
- which evidence supports which claim;
- which evidence contradicts another item;
- which source supersedes an older source;
- which measurement derives from which raw data;
- which conclusion depends on which intermediate inference.
This structure makes complex reasoning inspectable without flattening everything into a single confidence score.
74. Do not average away incompatible evidence
Suppose one source says 40, another says 80 and a third says 60.
The mean is 60. That arithmetic may be meaningless.
The sources may measure different quantities, periods or populations. One may be wrong. One may be a forecast and another an outcome. One may include tax while another excludes it.
Reconcile definitions before combining values.
75. Evidence review needs stop rules
Research can continue forever. Decisions cannot.
A stop rule asks when the evidence is sufficient for the purpose.
Possible stop conditions include:
- the decision is robust across remaining plausible uncertainty;
- new evidence is unlikely to change the conclusion materially;
- the cost of further evidence exceeds its expected decision value;
- a formal standard has been met;
- time constraints require a provisional decision with explicit uncertainty.
“Enough” is therefore purpose-sensitive.
76. Provisional conclusions are legitimate
Sometimes evidence supports action before it supports certainty.
A school may temporarily close a room after a safety concern while inspection continues. A researcher may report a preliminary estimate. An engineer may impose a conservative load limit while testing proceeds.
The correct label is not “fact established” if the evidence has not reached that level. It is “provisional action justified under current evidence and risk”.
77. Reversible decisions can tolerate different evidence thresholds
If a decision is cheap to reverse, it may be rational to act under less certainty and learn from the result.
If a decision is irreversible or potentially harmful, a higher threshold is usually appropriate.
Evidence is therefore part of an action architecture: uncertainty, consequence, reversibility and value of further information all matter.
78. For students: seven questions to ask of evidence
- What claim is this evidence supposed to support?
- How was the evidence produced?
- Is it about the right person, thing, time and place?
- What could make it wrong or misleading?
- Is another source genuinely independent?
- What evidence points the other way?
- How strong a conclusion does the evidence actually allow?
These questions work across English comprehension, Science, Mathematics, Humanities and everyday media reading.
79. For teachers: teach evidence as a relationship
Instead of asking only “Where is your evidence?”, ask:
- Which part of your answer does this evidence support?
- Why is it relevant?
- How reliable is it?
- What alternative explanation remains?
- What evidence would change your mind?
This moves students from quotation collection to reasoning.
80. For researchers: separate evidence generation from evidence interpretation
Study design creates observations under rules. Analysis transforms them. Interpretation connects results to claims. Generalisation extends those claims beyond the observed setting.
Each stage adds assumptions.
Transparent research makes those transitions visible instead of presenting the final conclusion as if it emerged directly from raw data.
For the wider methodology branch, use How Research Methods and Source Evaluation Work.
81. For editors: evidence should be auditable at sentence level
An editor should be able to pick a factual sentence and answer:
- what claim does this sentence make?
- what evidence supports it?
- does the evidence support every clause?
- is the source current?
- is there contradictory evidence?
- has the wording become stronger than the evidence?
Good prose should survive that inspection without becoming unreadable.
82. For AI systems: keep evidence separate from generated explanation
A safe public-facing reasoning pattern is:
- identify the claim;
- retrieve potentially relevant sources;
- resolve entity, time and scope;
- trace source lineage;
- extract the evidence that actually bears on the claim;
- record support and contradiction;
- assess uncertainty;
- form a bounded conclusion;
- retain the route back to source material.
The generated explanation should not become its own evidence merely because it is fluent.
83. A practical evidence state vocabulary
- supports — evidence bears in favour of the claim;
- strongly supports — relevant, reliable evidence substantially favours the claim within stated limits;
- contradicts — evidence bears against the claim;
- mixed — credible evidence points in different directions;
- insufficient — evidence exists but does not justify the proposed conclusion;
- not found — a search produced no qualifying evidence under its stated boundary;
- unknown — the evidential state does not support a conclusion;
- superseded — later evidence or a later authoritative state replaces an earlier current-state record;
- not relevant — the information does not bear materially on the claim.
These are practical communication labels, not a universal scientific grading standard.
84. The minimal rules
- Evidence is evidence for a claim.
- Data is not automatically evidence.
- A source can be strong for one claim and weak for another.
- A citation does not prove support.
- Relevance comes before prestige.
- Reliability depends on method and context.
- More copies are not more independent evidence.
- Provenance reveals lineage, not truth.
- Uncertainty belongs with measurement.
- Negative evidence requires coverage.
- Contradictory evidence stays visible.
- Freshness matters for changing claims.
- Evidence can be strong without being complete.
- Evidence can update belief without producing certainty.
- The conclusion must not be stronger than the evidence.
85. The final evidence checklist
- What exact claim is being evaluated?
- What kind of claim is it?
- What evidence bears directly on it?
- What is the evidence source?
- How was the evidence generated?
- Is the entity identity correct?
- Is the evidence from the relevant time?
- Does the population or scope match?
- What uncertainty does the measurement or method carry?
- Are apparently separate sources genuinely independent?
- What provenance connects them?
- What evidence contradicts the claim?
- What plausible alternative explanations remain?
- What is missing?
- Has any source been superseded?
- Would another observation materially change the decision?
- What conclusion is justified now?
- What conclusion would be too strong?
Conclusion: evidence is the disciplined connection between claim and world
Evidence is not the decoration around an argument. It is the route by which the argument touches reality.
A photograph, measurement, testimony, database, experiment, document or model output becomes useful evidence only when we know what claim it bears on and how the path from world to record should be interpreted.
That path can be strong or weak, direct or indirect, independent or copied, current or stale, precise or uncertain, complete or partial. Good reasoning does not hide those properties. It uses them.
The aim is not endless doubt. It is proportional confidence.
Believe no more than the evidence earns—but do not believe less merely because honest evidence includes uncertainty.
Sources and further reading
For philosophical background, see the Stanford Encyclopedia of Philosophy: Evidence. For provenance and derivation structures, see the W3C PROV-O Recommendation. For measurement uncertainty, see NIST Technical Note 1297 and NIST’s current Evidential Statistics programme. For evidence synthesis and certainty in intervention research, see the Cochrane Handbook, especially Chapter 14 on grading certainty of evidence. For an explicit domain example distinguishing sufficiency, relevance and reliability, see PCAOB AS 1105: Audit Evidence.
Within eduKateSingapore, continue with What Is a Claim?, What Is a Fact?, How Editorial Fact-Checking and Source Verification Work, How Research Methods and Source Evaluation Work, How Systematic Reviews and Evidence Synthesis Work, and How Archival Evidence Works.
eduKate Publishing. Public educational synthesis. Domain-specific standards are presented as examples within their own scope, not as universal evidence hierarchies. Fictional examples are used to teach the reasoning structure.
Research route: Return to the Research Collections Directory for connected guides on evidence, methods, discovery and scholarly systems.
