How to Categorise Documents and Archives | Type, Function, Provenance, Access and Retention

A document is more than a file, and an archive is more than a pile of old documents.

Documents carry content, purpose, authorship, format, rights, provenance, evidential value and lifecycle. Archives preserve relationships among records across time. Good categorisation therefore cannot rely on subject alone.

A contract, laboratory notebook, photograph, email, examination paper, invoice, manuscript, policy, map and database export may all refer to the same project while serving completely different functions.

Quick answer: how should documents and archives be categorised?

  • Document type: what form of record is it?
  • Function: what job did it perform?
  • Creator: who created or issued it?
  • Provenance: where did it come from and through whose custody?
  • Subject: what is it about?
  • Format: physical, digital, image, audio, video, structured data?
  • Lifecycle: draft, active, superseded, archived, disposed?
  • Access: public, internal, restricted, confidential?
  • Retention: how long must it be kept?
  • Evidence value: what claim or activity can it support?

This article applies the general framework from How to Categorise Anything to records and archival systems while connecting to the site’s specialist work on archival evidence.


1. Start with the record unit

Decide whether the unit is one file, one email, one folder, one case, one volume, one dataset or one collection. Mixing levels creates unstable categories.

2. Document type answers “what kind of record?”

Contract, report, invoice, memo, photograph and map are type labels. They describe form or documentary function, not necessarily subject.

3. Function often matters more than format

A PDF can be a contract, brochure, legal submission or teaching worksheet. File extension is not documentary meaning.

4. Creator should be explicit

Creator may be a person, office, company, committee, system or instrument. Creation responsibility is different from current ownership.

5. Provenance preserves origin

Records gain meaning from where they originated and the activity that produced them. Provenance can therefore be more informative than topic labels.

6. Custody is not authorship

An archive may hold a document it did not create. Separate creator, custodian, donor and current repository.

7. Fonds and series preserve context

Archival structures often classify records according to creator and activity before item-level subject. This protects relationships among records.

8. Subject classification remains useful

Subject supports discovery across creators and collections, but it should not erase provenance.

9. One document can have several subjects

Use multi-label or faceted subject indexing rather than forcing one dominant topic where several are genuinely present.

10. Format is an independent facet

Text, image, audio, video, spreadsheet, database, code and physical object require different handling and preservation strategies.

11. Carrier and content should be separated

The same content can move from paper to scan to OCR text. The intellectual record and its carrier are related but distinct entities.

12. Version is part of document identity

Draft, reviewed, approved, published and superseded versions may share content while having different authority.

13. Status should not be inferred from filename

“FINAL_v7_reallyfinal.pdf” is not governance. Store explicit lifecycle status and approval metadata.

14. Records have lifecycle states

Creation, active use, semi-active retention, archive and disposal describe different stages of record management.

15. Retention is not archival value

A record may be retained temporarily for operational or legal reasons without deserving permanent archival preservation.

16. Archival value can be evidential

A record may show how an institution acted, decided or communicated even if its subject content seems ordinary.

17. Archival value can be informational

Some records are preserved because the information they contain has long-term research or cultural value.

18. Access classification is separate

Public, internal, restricted and confidential should not be mixed into subject or type hierarchies.

19. Sensitivity can change over time

A record may be restricted during active use and later opened. Store effective dates and review rules.

20. Rights classification matters

Ownership, copyright, licence and permitted use are different legal dimensions.

21. Authenticity is not the same as accuracy

An authentic document can contain false statements. Classification should separate record authenticity from claim truth.

22. Originals and copies need relation types

Original, certified copy, scan, transcript, derivative and extract carry different evidential relationships.

23. Compound documents need structure

A report can contain appendices, charts, photographs and data tables. Preserve both whole-document and component relationships.

24. Email threads are relational records

Sender, recipient, reply sequence and attachments matter. Treating each email as isolated content loses conversation structure.

25. Case files are process containers

A case file may contain many document types connected by one transaction, student, patient, legal matter or project.

26. Dates need role labels

Creation date, publication date, effective date, receipt date and archival accession date are not interchangeable.

27. Location can be intellectual and physical

A record may belong intellectually to one collection while being stored physically or digitally elsewhere.

28. Digital records need technical metadata

File format, checksum, encoding, software dependency and migration history matter for preservation.

29. Digital preservation is not merely backup

Long-term preservation must keep content understandable and verifiable through format and system change.

30. OCR creates a derived representation

Recognised text should link back to the source image because transcription errors can affect search and evidence.

31. AI classification should preserve provenance

Models can assign subjects, document types and sensitivity candidates, but automated labels should record model version and confidence.

32. AI should not overwrite archival order blindly

Semantic clustering can improve discovery while original order and creator context remain intact.

33. Retention schedules require authority

Retention periods should identify the policy, regulation or institutional rule that governs them.

34. Disposal needs evidence

Records destroyed under an approved schedule should leave a disposal record showing what was removed, when and under which authority.

35. Duplicate detection is a classification problem

Exact duplicates, near-duplicates, revisions and derivatives should not be collapsed into one state automatically.

36. Archival uncertainty should be explicit

Unknown author, approximate date, disputed provenance and uncertain title are legitimate metadata states.

37. A practical document record

  • record ID;
  • document type;
  • function;
  • creator;
  • provenance;
  • subjects;
  • format;
  • version;
  • status;
  • access class;
  • rights;
  • retention rule;
  • evidence value;
  • related records;
  • confidence;
  • taxonomy version.

38. Classify the relationships, not only the items

Part-of, derived-from, supersedes, attached-to, created-by and transferred-to preserve the archival network.

39. Documents become archives through preserved context

Meaning grows when records remain connected to creators, activities, chronology and one another.

40. The deeper idea

A document is evidence of content and activity at the same time. Good categorisation preserves both.

The archive is not a shelf of isolated files. It is a memory system whose relationships are part of the evidence.

Final answer

Categorise documents and archives across type, function, creator, provenance, subject, format, lifecycle, access, rights, retention and evidential value. Preserve version and relationship structure, distinguish original from derivative, and keep archival context intact even when AI and semantic search add new discovery layers.


Continue through the series

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.