How Metadata Standards and Interoperability Work | From Dublin Core and MARC to DataCite, Schema.org, Crosswalks and Machine-Readable Meaning

Metadata is often described as “data about data”. The phrase is useful and incomplete. Metadata can describe a book, a dataset, a photograph, a person, a web page, a scientific sample, a software release, an archival file, a video, a museum object or a service. It can say what something is, who created it, when it was made, how it may be used, which version it is, where it belongs, how it relates to other things, what technical format it uses and how confidently those claims are known.

The problem is not creating metadata. Any organisation can invent fields. The difficult problem is creating metadata that another person or machine can interpret correctly later. That requires standards, identifiers, controlled vocabularies, schemas, namespaces, documented rules and governance.

Interoperability begins when two systems can exchange information. It becomes meaningful interoperability only when both systems interpret that information in sufficiently compatible ways. A field named date is not interoperable if one system means publication date and another means last-modified date. A field named author is not interoperable if one system stores a display string and another expects a persistent person identifier.

This article explains how metadata standards create the contract between data and meaning.

The metadata interoperability loop

RESOURCE
→ IDENTIFY
→ DESCRIBE
→ CHOOSE SCHEMA
→ DEFINE FIELDS
→ DEFINE VALUES
→ APPLY IDENTIFIERS
→ VALIDATE
→ SERIALISE
→ EXCHANGE
→ INTERPRET
→ CROSSWALK
→ PRESERVE PROVENANCE
→ UPDATE
→ REUSE

Failure can occur at every stage. A record can have the wrong identifier. A field can be populated with the wrong concept. A schema can be technically valid but semantically inappropriate. A crosswalk can lose distinctions. A serialisation can preserve syntax while dropping provenance. Interoperability is therefore a chain, not a file format.

1. Metadata answers questions about a resource

Metadata exists because resources need context. A digital image without metadata is pixels. Add creator, date, location, rights, subject and capture method, and the image becomes more discoverable and more usable as evidence.

Different metadata elements answer different questions:

A metadata scheme is therefore a theory of which questions matter for a particular resource and user community.

2. Metadata is not one layer

It is useful to distinguish several broad functions:

Metadata functionTypical jobExamples
DescriptiveDiscovery and identificationTitle, creator, subject, abstract
StructuralHow components fit togetherPage order, chapters, files in a package
AdministrativeManagement and controlOwner, access state, acquisition information
TechnicalHow a file or system is encodedFormat, dimensions, codec, checksum
RightsPermitted use and restrictionsLicence, copyright status, embargo
PreservationLong-term authenticity and usabilityEvents, agents, fixity, preservation actions
ProvenanceOrigin and transformation historySource, derivation, processing steps

The categories overlap. What matters is that metadata is designed around use cases rather than treated as a decorative description field.

3. A schema defines the shape of acceptable metadata

A metadata schema specifies elements, properties, structures or constraints used to describe resources. It may define names, meanings, datatypes, cardinalities, controlled values and relationships among fields.

A schema can answer questions such as:

Without these rules, two systems can use the same field name while storing fundamentally different structures.

4. Syntax and semantics are different

JSON, XML, CSV and RDF are ways of serialising information. They describe how data is encoded. A metadata standard describes what the encoded elements mean and how they should be used.

The same Dublin Core property can be expressed in RDF, XML, JSON or another compatible representation. Conversely, a perfectly valid JSON document can contain metadata whose field meanings are undocumented and therefore not interoperable.

File format answers “how is this written?” Metadata semantics answer “what does this mean?”

5. Namespaces prevent field-name collisions

Many schemas use common words such as title, creator, type and date. A namespace provides a globally distinguishable context so a property can be identified independently of its human-readable label.

In linked-data systems, a property is often identified by a URI. Two communities can both use a label “creator” while maintaining distinct formal definitions. A mapping can then state whether those properties are equivalent, broader, narrower or merely similar.

6. Dublin Core provides a compact cross-domain vocabulary

The Dublin Core Metadata Initiative maintains an authoritative set of metadata terms including the original fifteen-element Dublin Core set and a larger collection of properties, classes, datatypes and vocabulary encoding schemes.

Dublin Core became influential because it offers broadly reusable concepts such as title, creator, subject, description, publisher, contributor, date, type, format, identifier, source, language, relation, coverage and rights.

Its strength is cross-domain simplicity. Its weakness is the same thing: specialised communities often need richer rules than a generic element set alone can provide.

7. Application profiles turn general terms into local contracts

An application profile selects terms from one or more vocabularies and specifies how a particular community will use them. It can make some properties mandatory, restrict values, choose controlled vocabularies and define local rules without inventing an entirely new semantic universe.

This is often better than creating custom fields from scratch. The profile gains the interoperability of established terms while preserving the precision required by the local job.

8. MARC made library metadata machine-readable at scale

MARC—MAchine-Readable Cataloging—is a family of standards for representing and communicating bibliographic and related information in machine-readable form. The Library of Congress coordinates the MARC 21 formats with the wider cataloguing community.

MARC supports bibliographic, authority, holdings, classification and community information records. Fields and subfields encode detailed cataloguing structures developed over decades of library practice.

As of this edition, the official online MARC 21 documentation incorporates Update No. 42, May 2026. The Library of Congress notes that the full online format is the official standard and is updated as approved changes are integrated.

9. MARC’s longevity shows the value and cost of stable standards

MARC enabled enormous interoperability across library systems. It also carries structures shaped by earlier technological constraints and cataloguing traditions. Modern systems often need to map MARC into XML, linked data, discovery indexes or other models.

This is not evidence that the standard failed. It is evidence that successful standards accumulate history. The more systems depend on a standard, the more carefully change must be governed.

10. DataCite metadata makes research objects citable and discoverable

The DataCite Metadata Schema defines core properties for identifying, citing and discovering research data and other research outputs. It includes elements for identifiers, creators, titles, publisher, publication year, resource type, subjects, contributors, dates, related identifiers, rights and funding references.

Version 4.7, released 3 March 2026, added new controlled values including resource types such as Poster and Presentation and identifier relationships such as RAiD and SWHID support. The release history demonstrates schema governance in practice: metadata standards evolve while maintaining versioned documentation.

DataCite also shows how metadata and persistent identifiers reinforce one another. A DOI anchors the resource; the metadata makes the DOI discoverable and interpretable.

11. Schema.org brings structured metadata to the open web

Schema.org maintains a shared vocabulary for structured data on web pages, email and other contexts. It describes entities and relationships such as Article, Person, Organization, Event, Place, Course and many other types.

Schema.org can be encoded using JSON-LD, RDFa or Microdata. Search engines and other applications use this structured data to understand web content more explicitly than ordinary page text alone allows.

At the time of this edition, Schema.org displayed version 30.0, dated 19 March 2026. This again shows that public vocabularies are maintained editions rather than timeless dictionaries.

12. Structured data does not make a claim true

A web page can mark itself as an Article and name an author through Schema.org. That makes the claim machine-readable; it does not independently verify authorship. Metadata can faithfully encode false information.

Semantic structure improves interpretation. Evidence quality still depends on source authority, provenance and validation. See How Research Methods and Source Evaluation Work.

13. Controlled vocabularies stabilise metadata values

A schema can define a property called language, but interoperability remains weak if one system stores “English”, another “eng”, another “EN” and another “British English” without rules.

Controlled vocabularies and code lists reduce this variation. They may use ISO language codes, authority identifiers, resource-type vocabularies or domain-specific concept schemes.

This connects to How Authority Control and Controlled Vocabularies Work. Schemas define where values go; vocabularies help define what those values mean.

14. Identifiers stabilise the entities inside metadata

A creator field containing “Lee, J.” is difficult to disambiguate. A creator field carrying an ORCID iD can connect to a stable person identity. A related-resource field carrying a DOI can point precisely to a dataset or publication. An organisation identifier can distinguish institutions with similar names.

Persistent identifiers therefore turn metadata from descriptive text into a more reliable network of entity references. See How Persistent Identifiers Work.

15. Cardinality controls how many values a property can carry

A schema may allow one title, many titles, one publication date or many contributors. These rules matter. Flattening multiple creators into one comma-separated string destroys individual identity and makes later linking harder.

Cardinality is therefore semantic structure, not only validation syntax. It expresses what the model believes can meaningfully repeat.

16. Datatypes stop values from becoming ambiguous text

A date can be stored as a date datatype rather than a free-text note. A number can carry a numeric datatype. A URI can be explicitly identified as a URI. Language-tagged strings can preserve multilingual labels.

Typed values make validation and machine interpretation safer. They also expose errors earlier. The string “unknown” cannot be silently inserted into a field that is supposed to contain a number without triggering a design decision.

17. Required fields create minimum identity packets

Standards often designate some properties mandatory because a record without them cannot perform the intended job. DataCite, for example, requires core citation information around identifier, creator, title, publisher, publication year and resource type.

Required metadata should be the minimum sufficient for reliable operation. Requiring too little weakens discovery; requiring too much encourages low-quality filler or prevents valid resources from being registered.

18. Optional does not mean unimportant

Funding information, related identifiers, geographic coverage, methods and contributor roles may be optional in a general schema yet essential for a particular research question. Optionality reflects the general standard’s breadth, not the local importance of every field.

Application profiles can make globally optional fields locally mandatory where the collection requires them.

19. Metadata quality has several dimensions

A record can be schema-valid and still be wrong. Technical validation is necessary and insufficient.

20. Validation checks structure before interpretation

Machine validation can test whether required properties exist, whether dates match expected formats, whether values come from controlled lists and whether identifiers match syntactic rules.

Semantic validation goes further. Does the identifier actually resolve to the intended person? Is the publication year plausible? Is a dataset really a Dataset rather than Software? Does the licence apply to this version?

Strong metadata pipelines combine automated validation with domain-aware review.

21. Crosswalks translate between metadata schemes

A crosswalk maps elements in one metadata model to elements in another. A Dublin Core creator property might map to one or more creator structures in another schema. A MARC field may map into several linked-data properties.

Crosswalks are essential when systems exchange records without using identical schemas. They are also inherently risky because one model may distinguish concepts that the other merges.

22. Metadata mapping can be lossy

Suppose Schema A distinguishes creator, editor, translator and data curator, while Schema B has only contributor. Mapping A into B can collapse role distinctions. Mapping the simplified B record back into A cannot reconstruct the original roles.

RICH SOURCE MODEL
→ CROSSWALK
→ SIMPLER TARGET MODEL
→ INFORMATION LOSS
→ ROUND TRIP CANNOT RECREATE SOURCE

This is the metadata form of lossy compression. A crosswalk should document what cannot be preserved rather than pretending every field has a perfect equivalent.

23. Round-trip testing exposes hidden loss

One practical test is to map records from Schema A to Schema B and then back to A. Differences reveal information that the target model could not preserve.

Round-trip fidelity is not always required, but knowing the loss helps designers decide whether the crosswalk is appropriate for migration, discovery, archiving or only lightweight display.

24. Interoperability has several layers

LayerQuestion
TechnicalCan systems connect and exchange bytes?
SyntacticCan both parse the data structure?
SemanticDo both interpret fields and values compatibly?
OrganisationalDo governance, ownership and update processes align?
LegalMay the data be exchanged and reused?
TemporalCan versions and changes be reconciled through time?

A successful API call proves technical exchange. It does not prove semantic interoperability.

25. APIs carry metadata contracts into live systems

When metadata is exposed through an API, the schema becomes an operational contract. Clients depend on field names, meanings, datatypes and version behaviour.

Changing a field from string to object, renaming a controlled value or altering null semantics can break downstream systems even if the server remains online.

See Data APIs and Data Services for the wider service-contract layer.

26. Linked data represents metadata as relationships among identified things

Traditional records often look like self-contained forms. Linked data represents many metadata statements as relationships: resource A hasCreator person B; person B isAffiliatedWith organisation C; resource A isVersionOf resource D.

This graph model makes it easier to reuse shared identities instead of repeating descriptive strings in every record.

See Knowledge Graphs and Semantic Data for the broader architecture.

27. JSON-LD connects web-friendly JSON to linked-data semantics

JSON-LD allows developers to express linked-data concepts using JSON structures while connecting properties and values to globally identified vocabularies. Schema.org structured data is frequently embedded in web pages using JSON-LD.

The practical advantage is that ordinary web development can carry explicit semantic context without requiring every application developer to work directly with low-level RDF serialisation syntax.

28. Preservation metadata records what happened to digital objects

Long-term preservation requires more than title and creator. Institutions need to know what files existed, whether they changed, which preservation actions were performed, which software or agents were involved and whether fixity checks continued to pass.

The Library of Congress maintains resources for PREMIS, a widely used preservation metadata data dictionary. PREMIS models objects, events, agents and rights relevant to digital preservation.

This connects metadata directly to archival evidence and chain of custody.

29. Structural metadata lets compound objects remain intelligible

A digitised book can contain hundreds of page images, OCR files, a table of contents, cover images and supplementary material. Structural metadata tells a viewer which files belong together and in what order.

Without structure, the repository may preserve every file while losing the object-level experience. Preservation of parts is not automatically preservation of relationships among parts.

30. Technical metadata helps future systems render old files

Image dimensions, colour space, compression, codec, software version and file-format information can become essential when a future system needs to interpret a digital object. Technical metadata therefore supports migration, emulation, validation and quality control.

A file extension is not sufficient evidence of format. Strong preservation workflows may identify formats through signatures and record the tools used to inspect or transform them.

31. Rights metadata controls legitimate reuse

Open access is not the same as public visibility. A resource can be readable online while retaining restrictive copyright. A dataset can carry a specific licence. An archival record may have privacy or donor restrictions.

Rights metadata should identify the applicable licence or rights statement, relevant dates, restrictions and authority where known. Vague values such as “copyrighted” are often too weak for machine decisions.

32. Provenance metadata protects evidence through transformation

Data rarely remains untouched. Files are normalised, images resized, records merged, fields corrected, coordinates transformed and classifications updated. Provenance records how the current object was produced from earlier states.

Without provenance, a derived dataset may look authoritative while hiding a long chain of assumptions. With provenance, a researcher can trace the result back to sources, methods and software versions.

This connects to Data Versioning and Change Management.

33. Metadata itself has provenance

A title may come from the resource. A subject may be assigned by a librarian. A geographic coordinate may be generated by a geocoder. A description may be produced by AI and reviewed by a human. A rights statement may come from a donor agreement.

These metadata claims do not have equal evidential status. Mature systems can record who or what created a field, when, under which method and with what confidence.

34. Automatically generated metadata needs review rules

AI and machine-learning systems can extract entities, classify subjects, generate summaries, transcribe audio and identify objects in images. This can dramatically accelerate metadata creation.

Generated metadata should carry provenance and confidence because machine inference can be wrong. A face-recognition match, automated subject term or generated description is not equivalent to an authoritative human-verified record merely because both appear in the same field.

35. Metadata debt accumulates quietly

Systems often launch with minimal metadata and promise enrichment later. Later can become never. Missing identifiers, uncontrolled names, undocumented fields and weak rights information then make migration, discovery and AI retrieval increasingly expensive.

Metadata debt resembles technical debt: shortcuts create future maintenance costs. The difference is that lost provenance or historical context may become impossible to reconstruct after staff leave and source systems disappear.

36. Migration reveals whether semantics were documented

A system can function for years because staff informally know that status=3 means “approved”. During migration, undocumented conventions surface. The new system sees only a number.

Well-governed metadata externalises these meanings in schemas, code lists, data dictionaries and versioned documentation so knowledge survives the original developers.

37. Metadata standards need versioning

Schemas evolve because domains evolve. New resource types appear. identifiers become important. roles change. emerging practices need representation. Standards bodies release updates and document changes so implementations can migrate deliberately.

DataCite 4.7, MARC 21 Update 42 and Schema.org 30.0 are current examples of living metadata standards in 2026. Their version numbers are not clutter. They are part of the contract.

38. Backward compatibility is a policy choice

A standard can preserve old fields indefinitely, deprecate them, map them to replacements or make a breaking revision. Every choice has consequences for implementers and historical data.

Stable systems often prefer additive evolution because existing records and software are expensive to change. But permanent compatibility with every old design can also make a schema difficult to improve. Governance decides where to place the trade-off.

39. Cross-domain interoperability should preserve domain ownership

A medical record, a museum catalogue, a research dataset and a school resource should not be forced into one flat universal schema. Their domains require different detail and different authority structures.

Interoperability works better when shared core concepts—identifier, title, creator, date, rights, relation—are aligned while domain-specific extensions remain owned by the relevant community.

This is especially important inside eduKate: Biology, Medicine and Veterinary remain separate canonical owners even when shared metadata allows discovery across them.

40. Metadata can make a library machine-readable without making it machine-owned

Structured metadata helps search engines, APIs and AI systems navigate a library. But metadata models should still reflect human collection policy, evidence standards and editorial judgement. Machines consume the contract; they do not automatically define the collection’s purpose.

The stronger architecture is human-governed semantics with machine-readable expression.

41. A metadata design protocol

  1. Define the resource and reader job.
  2. Find an established community standard before inventing a local schema.
  3. Choose the minimum sufficient metadata core.
  4. Add a documented application profile for local requirements.
  5. Use persistent identifiers for entities where appropriate.
  6. Use controlled vocabularies for recurring categorical values.
  7. Define datatypes and cardinality.
  8. Record provenance and rights.
  9. Validate both syntax and meaning.
  10. Version the schema and publish change rules.

42. A crosswalk protocol

SOURCE SCHEMA
→ FIELD MEANING
→ TARGET CANDIDATE
→ EXACT / BROADER / NARROWER / NO MATCH
→ VALUE TRANSFORMATION
→ IDENTIFIER MAPPING
→ LOSS REGISTER
→ TEST RECORDS
→ ROUND-TRIP TEST WHERE RELEVANT
→ VERSION CROSSWALK
→ MONITOR

The loss register is important. A mapping project should say what cannot be translated, not only what can.

43. Metadata quality should be observable

Institutions can measure missing required fields, broken identifiers, invalid controlled values, stale links, duplicate entities, unlicensed records, unresolved provenance and crosswalk failures. Metadata should therefore participate in ordinary quality engineering.

See Data Quality and Data Testing and Reliability Engineering for the broader quality machinery.

44. The eduKate Library needs a metadata layer because prose alone cannot carry the whole estate

A human reader can infer that an article is about Singapore transport, was published in 2026 and relates to logistics. A machine should not have to infer every one of those relationships from prose on every visit. Structured metadata can make ownership, dates, types, identifiers, relationships and update states explicit.

This article therefore owns the general metadata-interoperability method. It does not replace the specific schemas of Science, Medicine, Veterinary, Publishing, Archives or any other collection. It explains how those owners can expose compatible shared metadata while retaining domain-specific meaning.

That is the correct direction for a large knowledge ecosystem: shared contracts where meanings align, explicit boundaries where they do not.

45. Metadata is institutional memory made portable

When metadata is good, a resource can move between systems without losing the explanation of what it is. A new repository can understand its identifier. A new librarian can understand its rights. A future program can distinguish creation date from modification date. An AI agent can follow explicit relationships rather than inventing them.

Portability is therefore not merely copying files. It is preserving the semantic contract around the files.

46. World Return from metadata standards

The World Return of metadata standards is reusable meaning. One institution describes a resource, and another can understand enough of that description to discover, cite, preserve, compare or connect it. A publisher’s article can enter a library catalogue. A dataset can enter a research graph. An archive can migrate systems. A web page can expose structured facts to search and AI.

The better the metadata contract, the less knowledge has to be rediscovered every time information crosses a boundary.

Sources and further reading

Continue through eduKate

Metadata is not decoration around information. It is the portable explanation that tells another system what the information is, how it is identified, how it relates to other things and what rules govern its interpretation. Standards make that explanation shareable. Interoperability begins when the explanation survives the boundary.

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading