How Citations, References and Scholarly Linking Work | From Claims and Sources to DOIs, Metadata, Provenance and the Research Graph

A citation is a tiny piece of text carrying an enormous responsibility. A surname and year, a footnote number or a DOI may be the only visible bridge between a claim and the evidence that supports it. If that bridge is weak, ambiguous or broken, a reader can no longer travel backwards from conclusion to source.

That is why citations are not decoration added after writing. They are part of research architecture. They help distinguish the author’s reasoning from borrowed evidence, make verification possible, connect present work to earlier work, reveal intellectual ancestry and allow later systems to build a graph of how knowledge developed.

The deeper principle is this: a scholarly claim should preserve a route back to the object that carries the evidence.

The citation chain

CLAIM
→ INLINE CITATION OR NOTE
→ REFERENCE ENTRY
→ PERSISTENT IDENTIFIER
→ METADATA RECORD
→ LANDING PAGE
→ VERSION / OBJECT
→ EVIDENCE
→ METHOD
→ WORLD

Not every source has every layer. A historical manuscript may have an archival identifier rather than a DOI. A law may have an official citation. A dataset may have a DataCite DOI. A standard may have a designation and issuing body. The general job remains the same: identify the source precisely enough that another reader can locate the same object and understand which version was used.

1. Citation, reference and bibliography are related but different

An in-text citation or note points from a specific passage to a source. A reference entry provides enough bibliographic information to identify that source. A reference list normally contains works actually cited. A bibliography can be broader and may include works consulted even if they are not directly cited.

The distinction matters because reference infrastructure should preserve the relationship between the exact claim and the exact source, not merely show that the author read widely.

2. The first job of a citation is attribution

Attribution tells the reader which ideas, data, words, methods or findings came from another source. This protects intellectual honesty and makes the author’s own contribution visible.

Attribution is especially important when paraphrasing. Changing the wording does not convert another researcher’s finding into the writer’s own discovery. The source relationship remains.

3. The second job is verification

A reader should be able to ask: does this source actually support this claim? Strong citation practice makes that check possible. Weak citation practice hides behind vague phrases such as “research shows” without telling the reader which research, what population, what result or what version.

Verification is one reason direct links to primary or authoritative sources are valuable when available. The citation becomes a route, not merely a badge of seriousness.

4. A source can support one claim and not another

A citation should be evaluated at claim level. A government report may be authoritative for an official statistic but not for a broad causal interpretation. A company may be authoritative for its own product specification but not for an independent comparison of competing products. A systematic review may support a pooled estimate while not supporting a claim about every subgroup.

This links citation practice directly to How Research Methods and Source Evaluation Work: source authority is local to the claim and method.

5. Citation styles are interfaces, not evidence hierarchies

APA, Chicago, MLA, Vancouver and other styles organise visible citation syntax differently. Some foreground author and date; others use notes or numbers. The style affects readability and disciplinary convention, but it does not determine source quality.

A perfectly formatted citation to a weak source remains weak. A slightly imperfectly formatted citation to an authoritative primary source may still preserve the evidential route, although publication standards should correct the formatting.

6. Bibliographic metadata makes references machine-readable

A human can often identify a work from author, title, journal and year. Machines need structured fields. Bibliographic metadata may include title, creators, publication date, publisher, container title, volume, issue, pages, identifiers, resource type, licence, funder, references and related objects.

Structured metadata allows search engines, libraries, repositories, citation managers and AI systems to distinguish one work from another and connect related records.

7. Persistent identifiers solve the problem of moving locations

A normal URL identifies a network location. If the publisher restructures its website, the URL may break. A persistent identifier separates identity from current location so that the identifier can continue resolving to the object’s present landing page.

The DOI system is the best-known scholarly example. A DOI does not guarantee quality, peer review or permanence of content. Its job is identity and resolution. The organisation responsible for the DOI must maintain the metadata and current destination.

See How Persistent Identifiers Work for the general identifier architecture.

8. Crossref turns references into research infrastructure

Crossref Reference Linking allows scholarly references to connect to persistent resources across publishers. Crossref also encourages members to deposit reference lists as metadata rather than leaving citations trapped only in article PDFs.

Its current documentation recommends supplying DOIs with references whenever possible and including full reference information when a DOI is not known. It also encourages citation of data, software and other supporting materials, not only journal articles.

This changes the scholarly record from a set of isolated documents into a machine-readable network of relationships.

9. Reference linking and depositing reference metadata are different operations

Crossref distinguishes linking a visible reference to a DOI from depositing the full reference list into Crossref metadata. Both are useful. The first helps readers navigate. The second makes citation relationships retrievable through metadata services and APIs.

This distinction matters because a website can display excellent references while contributing little machine-readable citation data, or deposit structured reference metadata while presenting references poorly to readers.

10. A DOI is stronger when accompanied by descriptive metadata

The identifier tells systems which object is meant. Metadata tells them what the object is. A DOI record with accurate title, creators, date, resource type, relationships and licence is far more useful than an identifier with sparse or stale metadata.

Crossref explains that its metadata is primarily supplied by members and includes relationships among research outputs, people, organisations, references, datasets and trials. That means metadata quality depends on the organisations maintaining those records.

11. DataCite extends citation infrastructure beyond conventional publications

DataCite supports persistent identification and metadata for research outputs such as datasets, software and other scholarly objects. Its documentation treats citations and references as explicit links among research outputs and encourages those relationships to be represented through DOI metadata.

This matters because modern research depends on far more than journal articles. A result may rely on a dataset, code repository, workflow, instrument package, model, preprint or protocol. Citation infrastructure should make those dependencies visible.

12. Data citation makes evidence lineage clearer

If an article analyses a dataset, the dataset should be identifiable as a first-class research object where possible. Citing only the article that originally introduced the dataset can hide which exact data version was used.

A useful data citation can identify creators, title, repository, year, version and persistent identifier. This helps later researchers reproduce analyses and gives credit to the work involved in producing reusable data.

13. Software citation preserves computational dependencies

Research increasingly depends on software. Different versions of a package can implement different algorithms, defaults or bug fixes. Merely writing “analysis was performed in software X” may not be enough to reproduce the result.

Where possible, software should be cited with version and persistent identity. The purpose is not ceremonial credit alone. It is to preserve the computational environment that contributed to the claim.

14. Versioning is central to honest citation

A living document, dataset, software package, standard or web page may change after it is cited. A reader needs to know which state supported the original claim.

Version-aware citation can include version numbers, publication dates, archive snapshots, edition identifiers or persistent links to specific releases. The correct method depends on the object type.

Citing the current homepage of a changing resource when the claim was based on an earlier version can silently rewrite evidence history.

15. Access date matters most when content can change

For stable published works, an access date may add little. For dynamic webpages, dashboards, live datasets and frequently revised guidance, recording when the content was consulted can be useful.

An access date does not preserve the content by itself. For high-stakes or historical verification, an archived version or versioned record is stronger.

16. Secondary citation can create evidence drift

Writers often encounter a claim in one source that cites another source. Repeating the claim while citing only the intermediary can gradually distort the original evidence, especially if each writer compresses the result further.

Whenever practical, important factual or quantitative claims should be checked against the closest available primary source. Secondary sources remain valuable for synthesis, interpretation and context, but they should not become accidental barriers between the claim and the evidence.

17. Citation chains can propagate errors

A paper can misquote an earlier result. Later papers can cite the misquotation. Eventually the claim appears well supported because many references repeat it, even though the chain traces back to one misunderstood source.

High citation count therefore does not automatically equal strong evidence. Citation networks measure attention and scholarly relationship, not truth.

18. Citation context matters

A cited work may be supported, criticised, replicated, contradicted or mentioned only as background. A raw citation count treats these relationships alike.

Machine-readable citation graphs become more informative when relationship type or citation context can be represented. The future of scholarly linking is therefore not only “A cites B” but increasingly “how and why A relates to B”.

19. Retractions and corrections change how a citation should be interpreted

A scholarly object can remain persistently identifiable even after correction or retraction. This is a feature, not a failure. The record should not disappear because later readers need to understand what happened.

Citation systems should therefore expose status changes. A DOI may still resolve, while Crossmark or publisher metadata communicates that the object has been corrected, updated or retracted.

The principle is familiar from archives: correction should preserve history rather than erase it.

20. Preprints and published articles are related but distinct objects

A preprint can later become a peer-reviewed publication, sometimes with substantial changes. A citation to the preprint should not be silently replaced in historical analysis by a citation to the later article unless the writer is explicitly updating the source relationship.

Metadata relationships can connect versions while preserving their distinct identities. This allows readers to follow the evolution of the work.

21. References should identify unusual object types explicitly

A standard, dataset, software package, conference paper, book chapter, report, law and journal article are different kinds of objects. Citation metadata becomes stronger when the object type is explicit rather than forcing everything into a journal-article template.

Crossref’s current reference-markup guidance supports citation types and specifically notes that data, software and standards can be represented in reference metadata. This helps machines avoid treating unlike objects as interchangeable.

22. Citation managers are transformation systems

Reference managers import metadata, store records and render them in chosen citation styles. They save enormous effort but can also preserve metadata errors indefinitely.

A DOI lookup may return an all-caps title, incomplete author list or inconsistent date. An imported webpage record may lack publisher information. Writers should therefore treat citation-manager metadata as a draft representation of the source, not automatically as authoritative truth.

23. Duplicate identifiers and duplicate records need reconciliation

The same work can appear in databases under variant titles, abbreviations or metadata states. Conversely, two different versions may look nearly identical. Reconciliation requires more than string matching.

Persistent identifiers, author identity, publication details, version relationships and provenance help determine whether records represent the same work, different manifestations or genuinely different objects.

24. Citations create a research graph

Once references are structured, every scholarly object becomes a node connected to other nodes. Articles cite articles. Articles cite data. Data can reference software. Corrections point to earlier versions. Reviews point to primary studies. Institutions, authors and funders connect through metadata.

This graph supports discovery, bibliometrics, systematic review searching, influence analysis and machine-assisted navigation. It also creates responsibility: incorrect links can propagate through many downstream systems.

25. Citation metrics are measurements with assumptions

Citation counts, h-indexes, journal metrics and related indicators can describe aspects of scholarly attention. They are not universal measures of truth, originality or social value.

Counts vary across disciplines, databases, publication ages and document types. Review articles often accumulate citations differently from datasets or methodological tools. Negative citations still count as citations. Any metric should therefore be interpreted inside its measurement model.

26. A citation network is incomplete by design

No citation database contains every scholarly relationship. Some publishers do not deposit full references. Some works lack persistent identifiers. Books, local journals, archival material and non-English scholarship may be unevenly represented.

Crossref itself explains that citation counts derived from its services are not comprehensive because not all scholarly works are registered and not all members deposit references. Missing links should therefore be treated as missing data, not proof that no relationship exists.

27. AI makes provenance more important, not less

Language models can generate plausible citations that do not exist, merge metadata from several real works, or cite a real source that does not support the claim. They can also help resolve references, locate DOIs and structure bibliographies.

The correct AI workflow is therefore:

CLAIM
→ AI SUGGESTS SOURCE CANDIDATES
→ VERIFY SOURCE EXISTS
→ RESOLVE IDENTIFIER
→ OPEN SOURCE
→ CHECK CLAIM SUPPORT
→ CHECK VERSION / STATUS
→ RECORD CITATION
→ PRESERVE PROVENANCE

AI can accelerate the path. It should not be allowed to invent the path.

28. Good citation practice is also good reading practice

A strong reader follows citations strategically. For a surprising claim, move to the source. For a review, inspect the primary studies behind an important conclusion. For a number, check definitions and denominator. For a standard, confirm edition. For a law, use the official text. For a historical statement, distinguish contemporary primary evidence from later interpretation.

Citation literacy is therefore a form of evidence literacy.

29. Citations in systematic reviews require study identity control

A systematic review can retrieve multiple publications from one underlying study. The references must allow those reports to be linked back to the study they describe. Trial registrations, persistent identifiers and study-level identifiers are therefore important parts of evidence synthesis.

See How Systematic Reviews and Evidence Synthesis Work for the full study-versus-report problem.

30. Citation infrastructure and scholarly publishing have different ownership

How Scholarly Publishing and Peer Review Work owns editorial submission, review and publication. This article owns the link architecture that connects published and unpublished research objects after and across publication.

How Metadata Standards and Interoperability Work owns the general problem of interoperable metadata. Citation infrastructure is one specific application where metadata relationships become evidential routes.

31. A practical citation audit

  1. Does the citation sit next to the claim it supports?
  2. Does the source actually support that claim?
  3. Is it the closest appropriate source?
  4. Is the source type clear?
  5. Is the version or edition correct?
  6. Is a persistent identifier available?
  7. Does the DOI resolve?
  8. Has the work been corrected or retracted?
  9. Are data and software dependencies cited where relevant?
  10. Can the reference be parsed by a human and a machine?
  11. Are secondary citations being mistaken for primary evidence?
  12. Could another researcher locate the same object?

32. Scholarly linking is memory infrastructure

Research advances because new work does not begin from zero. References allow a paper to inherit earlier methods, data, arguments and disagreements while showing where that inheritance came from.

Persistent identifiers and machine-readable metadata make that inheritance durable across websites and software systems. The result is larger than a bibliography. It is a distributed memory system for research.

Sources and authoritative infrastructure

Continue through eduKate

Wintour House return: A citation deserves the reader’s trust when it is more than a formatted token. It should preserve identity, version, relationship and provenance strongly enough that the reader can travel backwards from a sentence to the evidence, and future machines can reconstruct the same scholarly route without inventing missing links.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading