Metadata is often described as “data about data”. The phrase is useful and incomplete. Metadata can describe a book, a dataset, a photograph, a person, a web page, a scientific sample, a software release, an archival file, a video, a museum object or a service. It can say what something is, who created it, when it was made, how it may be used, which version it is, where it belongs, how it relates to other things, what technical format it uses and how confidently those claims are known.
The problem is not creating metadata. Any organisation can invent fields. The difficult problem is creating metadata that another person or machine can interpret correctly later. That requires standards, identifiers, controlled vocabularies, schemas, namespaces, documented rules and governance.
Interoperability begins when two systems can exchange information. It becomes meaningful interoperability only when both systems interpret that information in sufficiently compatible ways. A field named date is not interoperable if one system means publication date and another means last-modified date. A field named author is not interoperable if one system stores a display string and another expects a persistent person identifier.
This article explains how metadata standards create the contract between data and meaning.
The metadata interoperability loop
RESOURCE → IDENTIFY → DESCRIBE → CHOOSE SCHEMA → DEFINE FIELDS → DEFINE VALUES → APPLY IDENTIFIERS → VALIDATE → SERIALISE → EXCHANGE → INTERPRET → CROSSWALK → PRESERVE PROVENANCE → UPDATE → REUSE
Failure can occur at every stage. A record can have the wrong identifier. A field can be populated with the wrong concept. A schema can be technically valid but semantically inappropriate. A crosswalk can lose distinctions. A serialisation can preserve syntax while dropping provenance. Interoperability is therefore a chain, not a file format.
1. Metadata answers questions about a resource
Metadata exists because resources need context. A digital image without metadata is pixels. Add creator, date, location, rights, subject and capture method, and the image becomes more discoverable and more usable as evidence.
Different metadata elements answer different questions:
- What is it? Title, type, format, description.
- Who made it? Creator, contributor, publisher, institution.
- When? Created, issued, modified, valid, observed.
- Where? Spatial coverage, repository, location.
- About what? Subject, keywords, classification, abstract.
- Which one? Identifier, edition, version.
- Can I use it? Rights, licence, access restrictions.
- How was it produced? Method, instrument, software, provenance.
- How does it relate? IsPartOf, cites, replaces, versionOf, derivedFrom.
A metadata scheme is therefore a theory of which questions matter for a particular resource and user community.
2. Metadata is not one layer
It is useful to distinguish several broad functions:
| Metadata function | Typical job | Examples |
|---|---|---|
| Descriptive | Discovery and identification | Title, creator, subject, abstract |
| Structural | How components fit together | Page order, chapters, files in a package |
| Administrative | Management and control | Owner, access state, acquisition information |
| Technical | How a file or system is encoded | Format, dimensions, codec, checksum |
| Rights | Permitted use and restrictions | Licence, copyright status, embargo |
| Preservation | Long-term authenticity and usability | Events, agents, fixity, preservation actions |
| Provenance | Origin and transformation history | Source, derivation, processing steps |
The categories overlap. What matters is that metadata is designed around use cases rather than treated as a decorative description field.
3. A schema defines the shape of acceptable metadata
A metadata schema specifies elements, properties, structures or constraints used to describe resources. It may define names, meanings, datatypes, cardinalities, controlled values and relationships among fields.
A schema can answer questions such as:
- Is title required?
- Can there be more than one creator?
- Must a date follow a standard format?
- Is resource type chosen from a controlled list?
- Can a contributor carry an ORCID iD?
- Must a related resource identify the relationship type?
- Can values be language-tagged?
Without these rules, two systems can use the same field name while storing fundamentally different structures.
4. Syntax and semantics are different
JSON, XML, CSV and RDF are ways of serialising information. They describe how data is encoded. A metadata standard describes what the encoded elements mean and how they should be used.
The same Dublin Core property can be expressed in RDF, XML, JSON or another compatible representation. Conversely, a perfectly valid JSON document can contain metadata whose field meanings are undocumented and therefore not interoperable.
File format answers “how is this written?” Metadata semantics answer “what does this mean?”
5. Namespaces prevent field-name collisions
Many schemas use common words such as title, creator, type and date. A namespace provides a globally distinguishable context so a property can be identified independently of its human-readable label.
In linked-data systems, a property is often identified by a URI. Two communities can both use a label “creator” while maintaining distinct formal definitions. A mapping can then state whether those properties are equivalent, broader, narrower or merely similar.
6. Dublin Core provides a compact cross-domain vocabulary
The Dublin Core Metadata Initiative maintains an authoritative set of metadata terms including the original fifteen-element Dublin Core set and a larger collection of properties, classes, datatypes and vocabulary encoding schemes.
Dublin Core became influential because it offers broadly reusable concepts such as title, creator, subject, description, publisher, contributor, date, type, format, identifier, source, language, relation, coverage and rights.
Its strength is cross-domain simplicity. Its weakness is the same thing: specialised communities often need richer rules than a generic element set alone can provide.
7. Application profiles turn general terms into local contracts
An application profile selects terms from one or more vocabularies and specifies how a particular community will use them. It can make some properties mandatory, restrict values, choose controlled vocabularies and define local rules without inventing an entirely new semantic universe.
This is often better than creating custom fields from scratch. The profile gains the interoperability of established terms while preserving the precision required by the local job.
8. MARC made library metadata machine-readable at scale
MARC—MAchine-Readable Cataloging—is a family of standards for representing and communicating bibliographic and related information in machine-readable form. The Library of Congress coordinates the MARC 21 formats with the wider cataloguing community.
MARC supports bibliographic, authority, holdings, classification and community information records. Fields and subfields encode detailed cataloguing structures developed over decades of library practice.
As of this edition, the official online MARC 21 documentation incorporates Update No. 42, May 2026. The Library of Congress notes that the full online format is the official standard and is updated as approved changes are integrated.
9. MARC’s longevity shows the value and cost of stable standards
MARC enabled enormous interoperability across library systems. It also carries structures shaped by earlier technological constraints and cataloguing traditions. Modern systems often need to map MARC into XML, linked data, discovery indexes or other models.
This is not evidence that the standard failed. It is evidence that successful standards accumulate history. The more systems depend on a standard, the more carefully change must be governed.
10. DataCite metadata makes research objects citable and discoverable
The DataCite Metadata Schema defines core properties for identifying, citing and discovering research data and other research outputs. It includes elements for identifiers, creators, titles, publisher, publication year, resource type, subjects, contributors, dates, related identifiers, rights and funding references.
Version 4.7, released 3 March 2026, added new controlled values including resource types such as Poster and Presentation and identifier relationships such as RAiD and SWHID support. The release history demonstrates schema governance in practice: metadata standards evolve while maintaining versioned documentation.
DataCite also shows how metadata and persistent identifiers reinforce one another. A DOI anchors the resource; the metadata makes the DOI discoverable and interpretable.
11. Schema.org brings structured metadata to the open web
Schema.org maintains a shared vocabulary for structured data on web pages, email and other contexts. It describes entities and relationships such as Article, Person, Organization, Event, Place, Course and many other types.
Schema.org can be encoded using JSON-LD, RDFa or Microdata. Search engines and other applications use this structured data to understand web content more explicitly than ordinary page text alone allows.
At the time of this edition, Schema.org displayed version 30.0, dated 19 March 2026. This again shows that public vocabularies are maintained editions rather than timeless dictionaries.
12. Structured data does not make a claim true
A web page can mark itself as an Article and name an author through Schema.org. That makes the claim machine-readable; it does not independently verify authorship. Metadata can faithfully encode false information.
Semantic structure improves interpretation. Evidence quality still depends on source authority, provenance and validation. See How Research Methods and Source Evaluation Work.
13. Controlled vocabularies stabilise metadata values
A schema can define a property called language, but interoperability remains weak if one system stores “English”, another “eng”, another “EN” and another “British English” without rules.
Controlled vocabularies and code lists reduce this variation. They may use ISO language codes, authority identifiers, resource-type vocabularies or domain-specific concept schemes.
This connects to How Authority Control and Controlled Vocabularies Work. Schemas define where values go; vocabularies help define what those values mean.
14. Identifiers stabilise the entities inside metadata
A creator field containing “Lee, J.” is difficult to disambiguate. A creator field carrying an ORCID iD can connect to a stable person identity. A related-resource field carrying a DOI can point precisely to a dataset or publication. An organisation identifier can distinguish institutions with similar names.
Persistent identifiers therefore turn metadata from descriptive text into a more reliable network of entity references. See How Persistent Identifiers Work.
15. Cardinality controls how many values a property can carry
A schema may allow one title, many titles, one publication date or many contributors. These rules matter. Flattening multiple creators into one comma-separated string destroys individual identity and makes later linking harder.
Cardinality is therefore semantic structure, not only validation syntax. It expresses what the model believes can meaningfully repeat.
16. Datatypes stop values from becoming ambiguous text
A date can be stored as a date datatype rather than a free-text note. A number can carry a numeric datatype. A URI can be explicitly identified as a URI. Language-tagged strings can preserve multilingual labels.
Typed values make validation and machine interpretation safer. They also expose errors earlier. The string “unknown” cannot be silently inserted into a field that is supposed to contain a number without triggering a design decision.
17. Required fields create minimum identity packets
Standards often designate some properties mandatory because a record without them cannot perform the intended job. DataCite, for example, requires core citation information around identifier, creator, title, publisher, publication year and resource type.
Required metadata should be the minimum sufficient for reliable operation. Requiring too little weakens discovery; requiring too much encourages low-quality filler or prevents valid resources from being registered.
18. Optional does not mean unimportant
Funding information, related identifiers, geographic coverage, methods and contributor roles may be optional in a general schema yet essential for a particular research question. Optionality reflects the general standard’s breadth, not the local importance of every field.
Application profiles can make globally optional fields locally mandatory where the collection requires them.
19. Metadata quality has several dimensions
- Completeness: are required and useful fields populated?
- Accuracy: do values correctly describe the resource?
- Consistency: are rules applied uniformly?
- Validity: does the record conform to schema and datatype rules?
- Currency: are mutable facts up to date?
- Provenance: can users tell where metadata claims came from?
- Uniqueness: are duplicate records controlled?
- Interlinking: do identifiers and relationships connect correctly?
A record can be schema-valid and still be wrong. Technical validation is necessary and insufficient.
20. Validation checks structure before interpretation
Machine validation can test whether required properties exist, whether dates match expected formats, whether values come from controlled lists and whether identifiers match syntactic rules.
Semantic validation goes further. Does the identifier actually resolve to the intended person? Is the publication year plausible? Is a dataset really a Dataset rather than Software? Does the licence apply to this version?
Strong metadata pipelines combine automated validation with domain-aware review.
21. Crosswalks translate between metadata schemes
A crosswalk maps elements in one metadata model to elements in another. A Dublin Core creator property might map to one or more creator structures in another schema. A MARC field may map into several linked-data properties.
Crosswalks are essential when systems exchange records without using identical schemas. They are also inherently risky because one model may distinguish concepts that the other merges.
22. Metadata mapping can be lossy
Suppose Schema A distinguishes creator, editor, translator and data curator, while Schema B has only contributor. Mapping A into B can collapse role distinctions. Mapping the simplified B record back into A cannot reconstruct the original roles.
RICH SOURCE MODEL → CROSSWALK → SIMPLER TARGET MODEL → INFORMATION LOSS → ROUND TRIP CANNOT RECREATE SOURCE
This is the metadata form of lossy compression. A crosswalk should document what cannot be preserved rather than pretending every field has a perfect equivalent.
23. Round-trip testing exposes hidden loss
One practical test is to map records from Schema A to Schema B and then back to A. Differences reveal information that the target model could not preserve.
Round-trip fidelity is not always required, but knowing the loss helps designers decide whether the crosswalk is appropriate for migration, discovery, archiving or only lightweight display.
24. Interoperability has several layers
| Layer | Question |
|---|---|
| Technical | Can systems connect and exchange bytes? |
| Syntactic | Can both parse the data structure? |
| Semantic | Do both interpret fields and values compatibly? |
| Organisational | Do governance, ownership and update processes align? |
| Legal | May the data be exchanged and reused? |
| Temporal | Can versions and changes be reconciled through time? |
A successful API call proves technical exchange. It does not prove semantic interoperability.
25. APIs carry metadata contracts into live systems
When metadata is exposed through an API, the schema becomes an operational contract. Clients depend on field names, meanings, datatypes and version behaviour.
Changing a field from string to object, renaming a controlled value or altering null semantics can break downstream systems even if the server remains online.
See Data APIs and Data Services for the wider service-contract layer.
26. Linked data represents metadata as relationships among identified things
Traditional records often look like self-contained forms. Linked data represents many metadata statements as relationships: resource A hasCreator person B; person B isAffiliatedWith organisation C; resource A isVersionOf resource D.
This graph model makes it easier to reuse shared identities instead of repeating descriptive strings in every record.
See Knowledge Graphs and Semantic Data for the broader architecture.
27. JSON-LD connects web-friendly JSON to linked-data semantics
JSON-LD allows developers to express linked-data concepts using JSON structures while connecting properties and values to globally identified vocabularies. Schema.org structured data is frequently embedded in web pages using JSON-LD.
The practical advantage is that ordinary web development can carry explicit semantic context without requiring every application developer to work directly with low-level RDF serialisation syntax.
28. Preservation metadata records what happened to digital objects
Long-term preservation requires more than title and creator. Institutions need to know what files existed, whether they changed, which preservation actions were performed, which software or agents were involved and whether fixity checks continued to pass.
The Library of Congress maintains resources for PREMIS, a widely used preservation metadata data dictionary. PREMIS models objects, events, agents and rights relevant to digital preservation.
This connects metadata directly to archival evidence and chain of custody.
29. Structural metadata lets compound objects remain intelligible
A digitised book can contain hundreds of page images, OCR files, a table of contents, cover images and supplementary material. Structural metadata tells a viewer which files belong together and in what order.
Without structure, the repository may preserve every file while losing the object-level experience. Preservation of parts is not automatically preservation of relationships among parts.
30. Technical metadata helps future systems render old files
Image dimensions, colour space, compression, codec, software version and file-format information can become essential when a future system needs to interpret a digital object. Technical metadata therefore supports migration, emulation, validation and quality control.
A file extension is not sufficient evidence of format. Strong preservation workflows may identify formats through signatures and record the tools used to inspect or transform them.
31. Rights metadata controls legitimate reuse
Open access is not the same as public visibility. A resource can be readable online while retaining restrictive copyright. A dataset can carry a specific licence. An archival record may have privacy or donor restrictions.
Rights metadata should identify the applicable licence or rights statement, relevant dates, restrictions and authority where known. Vague values such as “copyrighted” are often too weak for machine decisions.
32. Provenance metadata protects evidence through transformation
Data rarely remains untouched. Files are normalised, images resized, records merged, fields corrected, coordinates transformed and classifications updated. Provenance records how the current object was produced from earlier states.
Without provenance, a derived dataset may look authoritative while hiding a long chain of assumptions. With provenance, a researcher can trace the result back to sources, methods and software versions.
This connects to Data Versioning and Change Management.
33. Metadata itself has provenance
A title may come from the resource. A subject may be assigned by a librarian. A geographic coordinate may be generated by a geocoder. A description may be produced by AI and reviewed by a human. A rights statement may come from a donor agreement.
These metadata claims do not have equal evidential status. Mature systems can record who or what created a field, when, under which method and with what confidence.
34. Automatically generated metadata needs review rules
AI and machine-learning systems can extract entities, classify subjects, generate summaries, transcribe audio and identify objects in images. This can dramatically accelerate metadata creation.
Generated metadata should carry provenance and confidence because machine inference can be wrong. A face-recognition match, automated subject term or generated description is not equivalent to an authoritative human-verified record merely because both appear in the same field.
35. Metadata debt accumulates quietly
Systems often launch with minimal metadata and promise enrichment later. Later can become never. Missing identifiers, uncontrolled names, undocumented fields and weak rights information then make migration, discovery and AI retrieval increasingly expensive.
Metadata debt resembles technical debt: shortcuts create future maintenance costs. The difference is that lost provenance or historical context may become impossible to reconstruct after staff leave and source systems disappear.
36. Migration reveals whether semantics were documented
A system can function for years because staff informally know that status=3 means “approved”. During migration, undocumented conventions surface. The new system sees only a number.
Well-governed metadata externalises these meanings in schemas, code lists, data dictionaries and versioned documentation so knowledge survives the original developers.
37. Metadata standards need versioning
Schemas evolve because domains evolve. New resource types appear. identifiers become important. roles change. emerging practices need representation. Standards bodies release updates and document changes so implementations can migrate deliberately.
DataCite 4.7, MARC 21 Update 42 and Schema.org 30.0 are current examples of living metadata standards in 2026. Their version numbers are not clutter. They are part of the contract.
38. Backward compatibility is a policy choice
A standard can preserve old fields indefinitely, deprecate them, map them to replacements or make a breaking revision. Every choice has consequences for implementers and historical data.
Stable systems often prefer additive evolution because existing records and software are expensive to change. But permanent compatibility with every old design can also make a schema difficult to improve. Governance decides where to place the trade-off.
39. Cross-domain interoperability should preserve domain ownership
A medical record, a museum catalogue, a research dataset and a school resource should not be forced into one flat universal schema. Their domains require different detail and different authority structures.
Interoperability works better when shared core concepts—identifier, title, creator, date, rights, relation—are aligned while domain-specific extensions remain owned by the relevant community.
This is especially important inside eduKate: Biology, Medicine and Veterinary remain separate canonical owners even when shared metadata allows discovery across them.
40. Metadata can make a library machine-readable without making it machine-owned
Structured metadata helps search engines, APIs and AI systems navigate a library. But metadata models should still reflect human collection policy, evidence standards and editorial judgement. Machines consume the contract; they do not automatically define the collection’s purpose.
The stronger architecture is human-governed semantics with machine-readable expression.
41. A metadata design protocol
- Define the resource and reader job.
- Find an established community standard before inventing a local schema.
- Choose the minimum sufficient metadata core.
- Add a documented application profile for local requirements.
- Use persistent identifiers for entities where appropriate.
- Use controlled vocabularies for recurring categorical values.
- Define datatypes and cardinality.
- Record provenance and rights.
- Validate both syntax and meaning.
- Version the schema and publish change rules.
42. A crosswalk protocol
SOURCE SCHEMA → FIELD MEANING → TARGET CANDIDATE → EXACT / BROADER / NARROWER / NO MATCH → VALUE TRANSFORMATION → IDENTIFIER MAPPING → LOSS REGISTER → TEST RECORDS → ROUND-TRIP TEST WHERE RELEVANT → VERSION CROSSWALK → MONITOR
The loss register is important. A mapping project should say what cannot be translated, not only what can.
43. Metadata quality should be observable
Institutions can measure missing required fields, broken identifiers, invalid controlled values, stale links, duplicate entities, unlicensed records, unresolved provenance and crosswalk failures. Metadata should therefore participate in ordinary quality engineering.
See Data Quality and Data Testing and Reliability Engineering for the broader quality machinery.
44. The eduKate Library needs a metadata layer because prose alone cannot carry the whole estate
A human reader can infer that an article is about Singapore transport, was published in 2026 and relates to logistics. A machine should not have to infer every one of those relationships from prose on every visit. Structured metadata can make ownership, dates, types, identifiers, relationships and update states explicit.
This article therefore owns the general metadata-interoperability method. It does not replace the specific schemas of Science, Medicine, Veterinary, Publishing, Archives or any other collection. It explains how those owners can expose compatible shared metadata while retaining domain-specific meaning.
That is the correct direction for a large knowledge ecosystem: shared contracts where meanings align, explicit boundaries where they do not.
45. Metadata is institutional memory made portable
When metadata is good, a resource can move between systems without losing the explanation of what it is. A new repository can understand its identifier. A new librarian can understand its rights. A future program can distinguish creation date from modification date. An AI agent can follow explicit relationships rather than inventing them.
Portability is therefore not merely copying files. It is preserving the semantic contract around the files.
46. World Return from metadata standards
The World Return of metadata standards is reusable meaning. One institution describes a resource, and another can understand enough of that description to discover, cite, preserve, compare or connect it. A publisher’s article can enter a library catalogue. A dataset can enter a research graph. An archive can migrate systems. A web page can expose structured facts to search and AI.
The better the metadata contract, the less knowledge has to be rediscovered every time information crosses a boundary.
Sources and further reading
- Dublin Core Metadata Initiative — DCMI Metadata Terms
- Library of Congress — MARC Standards
- Library of Congress — MARC 21 official format status and updates
- DataCite — Metadata Schema 4.7
- Schema.org — Structured data vocabulary
- Library of Congress — PREMIS Preservation Metadata
- W3C — JSON-LD 1.1
- Library of Congress — Linked Data Service
Continue through eduKate
- How Libraries Work
- How Persistent Identifiers Work
- How Authority Control and Controlled Vocabularies Work
- How Standards Work
- Knowledge Graphs and Semantic Data
- How Data Management Works
- Data Versioning and Change Management
- Data Testing and Reliability Engineering
Metadata is not decoration around information. It is the portable explanation that tells another system what the information is, how it is identified, how it relates to other things and what rules govern its interpretation. Standards make that explanation shareable. Interoperability begins when the explanation survives the boundary.
