To categorise anything well, do not begin with boxes. Begin with the question the boxes are supposed to help you answer.
We categorise constantly. A child sorts animals. A doctor classifies a disease. A librarian assigns a book. A historian divides time into periods. A business groups customers. A scientist names species. A government defines industries. A museum describes an object. A search engine decides what a page is about. An artificial-intelligence system maps a prompt to a task, a tool, a domain and a route.
The surface problem looks simple: put the thing in the right category.
The real problem is harder. What is the thing? Which features matter? Who is asking? For what purpose? At what scale? During which period? Can it belong to more than one category? What happens when it changes? What happens when the evidence is incomplete? What happens when two careful people classify it differently?
This is why a good classification system is not merely a list of labels. It is a controlled way of turning messy reality into a useful representation without pretending the representation is reality itself.
Quick answer: how do you categorise anything?
- Define the object. State exactly what is being classified.
- Define the purpose. Decide what the classification must help someone do.
- Choose the unit. Do not mix people, events, organisations, documents, properties and processes in one undifferentiated layer.
- Choose dimensions. Ask which independent axes matter: form, function, origin, behaviour, time, scale, risk, use, ownership, evidence or another relevant dimension.
- Write criteria. Each category needs an admission rule, not merely a memorable name.
- Allow overlap when reality overlaps. Use facets, tags or multiple inheritance instead of forcing false exclusivity.
- Record uncertainty. Unknown, disputed, provisional and not-applicable are legitimate states.
- Test edge cases. A category system is only as strong as the awkward examples it can explain.
- Separate description from judgement. Classification should not quietly turn into ranking or moral scoring.
- Preserve provenance. Record who classified the item, when, from which evidence and under which version of the scheme.
- Reclassify when the world changes. Categories are tools; reality is allowed to move.
- Measure usefulness. The best system is not the one with the most categories. It is the one that improves retrieval, comparison, reasoning, prediction or action.
In Singapore English, this article uses categorise. The American spelling is categorize. The method is the same.
1. Categorisation is a model of the world
A category is not the thing itself. It is a statement about the thing.
A whale exists whether humans call it a fish, a mammal, a marine animal, a vertebrate, an endangered species, a cetacean, a tourist attraction, a conservation concern or a biological population. Each label selects a different relationship between the whale and the question being asked.
This gives us the first rule:
Categories do not merely divide reality. They reveal which distinctions we have chosen to care about.
That choice may be sensible, arbitrary, historical, political, scientific, administrative or practical. Good categorisation makes the choice visible.
2. Start with the job, not the labels
Imagine a box of objects: screwdriver, spoon, battery, passport, apple, medicine bottle, USB cable, notebook and key.
There is no single correct way to categorise them. A household might sort by room. A shop sorts by product class. An airport sorts by security rules. A recycling centre sorts by material. An archive might care about ownership and date. A child may sort by colour. An insurer may sort by risk. A logistics company may sort by size, fragility and storage conditions.
The same object can therefore receive several correct classifications because the job has changed.
Before building categories, finish this sentence:
We are classifying these things so that we can ______.
Retrieve them faster? Compare them? Teach them? Regulate them? Diagnose them? Route them? Count them? Store them? Predict behaviour? Detect change? Allocate resources? Decide responsibility?
If the blank is unclear, the taxonomy will eventually become unclear too.
3. Define the object before you classify it
Many classification failures begin one step too late. People debate categories before agreeing on what is being categorised.
Are we classifying a physical object, an event, a person, an organisation, a place, a process, a claim, a document, a state, a relationship, a role or an idea?
A hospital is an organisation, a building, a legal entity, an employer, a service system and a place. A book is a text, an edition, a physical copy, an intellectual work and an archival object. A company can be a registered entity, a brand, a workforce, a set of contracts, a group of assets or an operating system.
If these levels are mixed, categories begin to collide. One branch describes what something is; another describes what it does; another describes where it is located; another describes who owns it. They look like siblings but they are different kinds of statements.
4. Give the object a boundary
Every classification needs a boundary. Otherwise the object leaks into its surroundings.
If we classify a school, do we include the building, staff, students, curriculum, alumni, parent community, contractors, online systems and governing organisation? If we classify a city, do we mean administrative boundaries, the continuous built-up area, the labour market, the transport region or the cultural idea of the city?
Write three coordinates before classification begins:
- Where? Spatial boundary.
- When? Temporal boundary.
- What is included? System boundary.
The same lesson appears in the eduKateSingapore article How to Categorise Civilisation?: a civilisation is usually better represented as a set of coordinates than forced into one permanent box.
5. Distinguish identity from properties
One of the most useful moves in classification is to separate what something is from what is true about it.
A car may be an electric passenger vehicle. It may also be red, four years old, leased, damaged, registered in Singapore, parked indoors and used for ride-hailing. The first statement describes a relatively stable class. The others are properties, states, relationships or uses.
If every property becomes a category branch, the tree explodes. If no properties are recorded, the classification becomes too coarse.
A practical representation often needs both:
- Type: what kind of thing is this?
- Attributes: what characteristics does it have?
- State: what condition is it in now?
- Relationships: how is it connected to other things?
- Context: where, when and for whom is the statement true?
6. Choose dimensions before categories
A powerful classification begins with dimensions. A dimension is an axis along which things can differ.
Common dimensions include:
- Form — what shape or structure does it have?
- Function — what does it do?
- Origin — where did it come from?
- Mechanism — how does it work?
- Material — what is it made from?
- Scale — how large, numerous or extensive is it?
- Time — when did it exist or occur?
- Location — where is it?
- Ownership — who controls or possesses it?
- Use — how is it being used?
- Risk — what can go wrong?
- Capability — what can it reliably do?
- Evidence — how certain are we?
- Lifecycle — created, active, dormant, retired, archived, destroyed?
When dimensions are explicit, classification becomes composable. A library can describe a document by subject, author, language, date, format, audience and rights without pretending these are competing answers to one question.
7. One tree is rarely enough
The simplest taxonomy is a tree: one root divides into branches, each branch divides again, and every item descends through one path.
Trees are excellent when the world genuinely has nested containment or when the operating job requires one authoritative route. They are less effective when objects participate in several legitimate structures simultaneously.
Consider a book about the history of medicine in colonial Singapore written in English for university students. Is it history, medicine, Singapore studies, colonial studies, English-language material, tertiary education or social history?
Yes.
The correct design may be a faceted system: subject, place, time, audience, language and format remain separate dimensions that can be combined during retrieval.
8. Use hierarchy when the relationship is really “is a”
Hierarchies are useful, but only when each step preserves the same kind of relationship.
A sparrow is a bird. A bird is a vertebrate animal. That is a type hierarchy.
A wing is part of a bird. That is not the same relationship. A bird lives in a habitat. A bird eats food. A bird may be protected by law. A bird may be observed in Singapore. These are relationships, not subclasses.
A classification system becomes much clearer when it does not disguise every relationship as parent and child.
9. Write admission rules
A category name without criteria is only a suggestion.
For each important category, write:
- Definition: what does this category mean?
- Required criteria: what must be true?
- Exclusion criteria: what disqualifies an item?
- Evidence rule: what evidence is sufficient?
- Boundary examples: which difficult cases illustrate the edge?
- Counterexamples: what looks similar but does not belong?
This is the difference between saying “large company” and stating what large means for the purpose at hand. Employee count? Revenue? Market capitalisation? Assets? Geographic footprint? A regulator, investor and economist may each need a different threshold.
10. Categories need negative space
A useful definition does not only tell you what belongs. It tells you what does not.
If a learner cannot identify the nearest counterexample, the category may still be fuzzy. In science education this is especially important. The site’s guide Understanding Why Classification Needs Clear Criteria treats classification as evidence-based comparison rather than decorative labelling.
Ask three questions:
- What is the clearest member of this category?
- What is the nearest non-member?
- What is the case that makes us hesitate?
Those three examples reveal more about the real boundary than ten easy examples.
11. Not every category has a sharp boundary
Some categories are defined by strict rules. A legal age threshold may be exact. A chemical element is defined by proton number. A calendar date is discrete.
Other categories behave more like clusters. “Game”, “furniture”, “genre”, “city”, “intelligence”, “friend”, “luxury”, “emergency” and “culture” can have central examples without one simple feature shared by every member.
For these categories, a prototype model may be more honest: some members are highly typical, others are peripheral, and membership may depend on context.
Do not force a crisp boundary merely because the database would prefer one.
12. Allow an item to belong to more than one category
False exclusivity is one of the most common design mistakes.
A person can be a parent, engineer, citizen, customer, volunteer, patient, alumnus and traveller. A photograph can be historical evidence, an artwork, a family record, a legal exhibit and a teaching resource. A river can be an ecosystem, water source, border, transport corridor, sacred place and flood hazard.
When reality carries several roles, the representation should be able to carry several roles too.
13. Separate permanent type from temporary state
A school can be open, closed for holidays, under renovation or permanently closed. A machine can be active, idle, faulty, repaired or retired. A patient can be stable, deteriorating, recovering or discharged.
These are states, not necessarily new types of object.
State deserves its own field because state changes quickly. If state is baked into the taxonomy, every change requires structural reclassification and the history becomes hard to reconstruct.
14. Time is not metadata you can forget
Many classification statements are only true during a period.
A company changes industry. A territory changes political status. A document moves from confidential to public. A disease definition changes. A planet may be reclassified. A neighbourhood changes land use. A language evolves. A technology moves from experimental to ordinary.
Record valid from and, where relevant, valid until. The question “What category is this?” often needs to become “What category was this at that time?”
15. Context can change the correct answer
A tomato is botanically a fruit and culinarily treated as a vegetable. That familiar example is not a joke about classification. It demonstrates that a word may sit inside different systems built for different jobs.
Contextual classification is legitimate when the context is explicit. Trouble begins when one context silently borrows the authority of another.
Write the viewpoint into the system when needed: scientific, legal, commercial, educational, archival, operational, cultural or colloquial.
16. Unknown is a category of knowledge, not a failure
Classification systems often make uncertainty disappear because software forms prefer a completed field.
That creates false certainty.
A strong system distinguishes:
- Known and classified
- Known but not yet classified
- Unknown
- Uncertain
- Disputed
- Not applicable
- Insufficient evidence
- Outside current scope
These states protect the integrity of the dataset. “I do not know yet” is often more valuable than a confident wrong label.
17. Do not let “Other” become a warehouse
Most taxonomies eventually need an “Other” category. That is acceptable as a pressure-release valve. It becomes dangerous when “Other” grows faster than the named categories.
A large “Other” bucket is a sensor. It may mean the taxonomy is outdated, the dimensions are wrong, a new category has emerged, or users are applying the rules inconsistently.
Review “Other” periodically. Its contents are often the future structure trying to become visible.
18. The right number of categories is a design decision
Too few categories make everything look the same. Too many make every item unique and destroy the benefit of grouping.
The useful level depends on the job. A school science exercise may need “mammal, bird, fish, reptile, amphibian”. A zoological research database needs much finer resolution. A supermarket may care about “dairy” while a dietitian cares about lactose content, protein source, processing and allergens.
Resolution should therefore be fit for purpose. Do not confuse precision with usefulness.
19. Good categories should be discriminating
A useful category changes what you can do next.
If two categories produce exactly the same retrieval, treatment, explanation, storage rule, risk response or decision, ask whether the distinction is operationally useful.
This does not mean every intellectually meaningful difference needs an immediate action. It means the architecture should know why the distinction exists.
20. Classification is compression
Every object has more detail than a category can preserve. Categorisation compresses information.
Calling something “a bridge” ignores its exact dimensions, material, age, load rating, design history, ownership, condition, traffic volume and cultural significance. The label is useful because it discards detail while preserving a pattern relevant to many tasks.
Good compression keeps the information needed for the next step. Bad compression throws away the distinction that later turns out to matter.
21. Classification can create blind spots
Once categories exist, people begin seeing through them. Reports count what has a code. Budgets follow recognised programmes. search systems retrieve indexed topics. Institutions respond to named risks.
Anything that falls between categories can become less visible.
This is why every mature classification system needs a mechanism for anomalies, boundary cases and new phenomena. A taxonomy should not punish reality for failing to resemble its schema.
22. Separate categories from rankings
Classification answers “what kind?” Ranking answers “how much?” or “which is better?” These are different operations.
Advanced/basic, developed/developing, high/low, premium/ordinary and normal/abnormal can quietly mix description with judgement. Sometimes a threshold is required. Sometimes the label carries assumptions that should be made explicit.
A city can be large without being well governed. A school can be selective without being effective for every student. A technology can be newer without being better for a particular task. A civilisation can be technologically sophisticated and institutionally brittle at the same time.
Do not collapse multiple dimensions into one ladder unless the job genuinely requires a scalar score.
23. Separate description from moral judgement
A classification can be descriptively accurate without approving what it describes. Conversely, moral concern does not remove the need for accurate description.
This matters in history, politics, medicine, law, social science and any field where categories affect people. A system should state whether a label describes structure, behaviour, risk, legal status, identity, diagnosis, administrative treatment or ethical evaluation.
When these layers are confused, labels become arguments disguised as facts.
24. Preserve the evidence behind the label
A category without provenance is difficult to audit.
For important decisions, store:
- the classification;
- the evidence used;
- the rule or version applied;
- the classifier or source;
- the date;
- the confidence or certainty level;
- any competing classification;
- the reason for later changes.
This turns classification from a silent label into an auditable claim.
25. Test the taxonomy before trusting it
Do not validate a category system only with the examples used to invent it.
Test at least six classes of cases:
- Obvious members — the easy cases.
- Obvious non-members — the nearest clear contrast.
- Boundary cases — examples near the definition edge.
- Hybrids — items with multiple legitimate memberships.
- Novel cases — things that did not exist when the taxonomy was written.
- Adversarial cases — examples that exploit ambiguous wording or conflicting rules.
If the scheme only works on easy examples, it is a vocabulary list rather than an operating classification system.
26. Measure agreement
Give the same sample to several independent classifiers. If they consistently disagree, the problem may not be the people. The definitions may be underspecified.
Disagreement is diagnostic. Ask which categories attract inconsistent decisions, which criteria are interpreted differently and whether the evidence needed for classification is actually available.
The goal is not to eliminate every disagreement. Some domains are inherently uncertain. The goal is to know where and why disagreement occurs.
27. Version the classification system
Taxonomies change. Categories are renamed, split, merged, retired or given new definitions.
Do not overwrite history as if the old scheme never existed. Keep versions and crosswalks. A record classified under Version 2 should still be interpretable after Version 5 arrives.
A useful change log records:
- old category;
- new category or categories;
- relationship: same, narrower, broader, split, merge or retired;
- effective date;
- reason for change;
- whether old records should be migrated.
Crosswalking is how a classification system remembers its own past.
28. Categories should support retrieval
One of the oldest reasons to classify is simple: find the thing again.
Libraries, archives, warehouses, hospitals, museums, websites and databases all face the same problem at different scales. Storage without retrieval is accumulation, not organised knowledge.
A strong classification system improves both:
- precision — fewer irrelevant results;
- recall — fewer relevant things missed.
Sometimes the best architecture combines taxonomy with search, metadata, tags and relationships rather than expecting the taxonomy to carry every retrieval job alone.
29. Categories should support comparison
Classification becomes especially powerful when the same dimensions are applied consistently across many objects.
Once every city is described by population, area, transport, energy, water, governance, economy and risk, comparison becomes possible. Once every document records author, date, subject, format, provenance and rights, an archive can see patterns across the collection. Once every medical case records compatible variables, outcomes can be compared.
The category system becomes a measurement frame.
30. Categories should support action without pretending to be action
A label can trigger a workflow: hazardous material, urgent case, archival restriction, premium customer, endangered species, high-priority fault.
That can be useful, but the label and the action should remain distinguishable. If the action changes, the underlying classification may remain valid. If the classification is uncertain, the action may require a precautionary rule.
Think in three layers:
- Observation: what do we know?
- Classification: what category best represents that evidence?
- Decision: what should we do?
Keeping these layers separate makes errors easier to find and decisions easier to revise.
31. A universal categorisation worksheet
For almost any domain, answer the following in order:
- What exactly is the object?
- Why are we categorising it?
- Who will use the result?
- What decisions or retrieval tasks must it support?
- What is the object boundary?
- What is the time boundary?
- Which dimensions matter?
- Which dimensions should stay separate?
- Which relationships are “is a”, “part of”, “located in”, “used for”, “owned by” or another relation?
- Which categories need formal criteria?
- What are the nearest counterexamples?
- Can items legitimately have multiple memberships?
- What states can change over time?
- How will uncertainty be represented?
- How will new cases be handled?
- How will categories be tested?
- How will disagreements be reviewed?
- How will versions and crosswalks be stored?
- How will the system be measured?
- What would make us retire or redesign the taxonomy?
32. The four tests of a useful category
Before keeping a category, run four tests.
The meaning test
Can two trained users explain the category in roughly the same way?
The boundary test
Can they explain the nearest member, non-member and ambiguous case?
The utility test
Does the distinction improve retrieval, comparison, explanation, prediction, allocation or action?
The survival test
Can the category absorb new examples without constant emergency repair?
A category that fails all four tests may be a word in search of a job.
33. From categorisation to a knowledge architecture
Once classification becomes large, it stops being only a taxonomy problem. It becomes knowledge architecture.
You may need:
- taxonomy for type hierarchies;
- facets for independent dimensions;
- metadata for descriptive attributes;
- relationships for connections between entities;
- controlled vocabulary for consistent terms;
- identifiers for stable reference;
- provenance for evidence and authorship;
- versioning for change through time;
- crosswalks for mapping between systems;
- search for flexible retrieval;
- human review for cases the schema cannot safely settle.
This is the larger purpose of the World Knowledge Research Library Projection: not merely storing knowledge, but making different bodies of knowledge findable, comparable and traversable.
34. What artificial intelligence changes
AI makes classification faster, but it does not remove the need for a schema.
A model can propose labels, infer similarity, cluster documents, extract entities and identify candidate relationships. But the system still needs to know the purpose of classification, the authority of sources, the allowed categories, the acceptable uncertainty and the action consequences of a label.
AI also makes one design principle more important: keep the evidence and the reasoning path inspectable. When a machine classifies at scale, silent category errors can propagate at scale too.
The best architecture therefore combines machine speed with explicit definitions, provenance, confidence, exception handling and review.
35. The deepest rule: categorise relationships, not only things
Many difficult classification problems become easier when we stop asking only “What is this?” and start asking “How is this connected?”
A teacher teaches a student. A medicine treats a condition. A road connects places. A document supports a claim. A company owns an asset. A species inhabits an ecosystem. A law applies within a jurisdiction. A historical event causes consequences. An archive preserves a record.
Once relationships are explicit, one object no longer needs to carry every fact inside one giant label. Knowledge can become a network rather than a filing cabinet.
36. A category is a promise
When a system assigns a label, it makes a promise to the next user: items sharing this label are similar in a way that matters for your job.
If that promise is false, the category becomes noise. If it is partly true but the conditions are hidden, the category becomes misleading. If the promise is clear and consistently kept, the category becomes infrastructure.
This is why classification can look administrative while being intellectually deep. It sits between reality and every system that tries to understand reality.
Final answer
How do you categorise anything?
Define the object. Define the job. Give the object coordinates. Separate type from property, state and relationship. Choose dimensions before labels. Write criteria. Allow overlap. Preserve uncertainty. Test boundaries. Record provenance. Version the scheme. Measure whether the categories actually help someone find, compare, understand or act.
And remember the final constraint: the map is allowed to be neat; the world is not required to be.
Continue through the How to Categorise Anything series
- How Categories Work | Boundaries, Similarity, Prototypes and Exceptions
- How to Build a Taxonomy | From Messy Reality to a Usable Classification System
- How to Test a Classification System | Edge Cases, Drift, Error and Revision
- How to Categorise Civilisation?
- World Knowledge Research Library Projection
More articles in this collection
Classification guides: roles, decisions and systems
- How to Categorise Access | Public, Internal, Restricted, Conditional, Temporary and Privileged
- How to Categorise Arguments | Claim, Premises, Inference, Evidence, Objections and Strength
- How to Categorise Authority | Source, Scope, Delegation, Jurisdiction, Duration and Review
- How to Categorise Boundaries | Physical, Conceptual, Administrative, Temporal, Permeable and Contested
- How to Categorise Capabilities | Function, Maturity, Capacity, Readiness, Dependency and Evidence
- How to Categorise Changes | Type, Direction, Magnitude, Rate, Cause, Reversibility and Consequence
- How to Categorise Commitments | Promise, Contract, Plan, Deadline, Dependency, Evidence and Status
- How to Categorise Controls | Preventive, Detective, Corrective, Manual, Automated and Compensating
- How to Categorise Estimates | Quantity, Method, Range, Confidence, Assumptions, Sensitivity and Revision
- How to Categorise Feedback | Informational, Evaluative, Corrective, Delayed, Continuous and Closed-Loop
- How to Categorise Incentives | Positive, Negative, Intrinsic, Extrinsic, Immediate, Delayed and Perverse
- How to Categorise Instructions | Goal, Preconditions, Sequence, Branching, Constraints, Exceptions and Verification
- How to Categorise Intentions | Goal, Specificity, Commitment, Time Horizon, Conditions, Evidence and Revision
- How to Categorise Interfaces | Boundary, Contract, Input, Output, Protocol, Compatibility and Failure
- How to Categorise Interpretations | Object, Context, Framework, Evidence, Ambiguity and Alternatives
- How to Categorise Labels | Descriptive, Status, Sensitivity, Workflow, Compliance and Machine-Readable
- How to Categorise Obligations | Source, Subject, Action, Deadline, Condition, Enforcement and Discharge
- How to Categorise Ownership | Legal, Beneficial, Custodial, Shared, Temporary and Contested
- How to Categorise Permissions | Subject, Action, Resource, Scope, Condition, Duration and Revocation
- How to Categorise Plans | Goal, Sequence, Resources, Dependencies, Milestones, Contingencies and Review
- How to Categorise Policies | Purpose, Authority, Scope, Rules, Exceptions, Enforcement and Version
- How to Categorise Preferences | Object, Strength, Ranking, Stability, Context, Trade-Offs and Evidence
- How to Categorise Properties | Intrinsic, Relational, Measured, Derived, Stable and Dynamic
- How to Categorise Requirements | Source, Necessity, Scope, Priority, Verification and Change
- How to Categorise Responsibilities | Owner, Scope, Authority, Duty, Evidence, Escalation and Accountability
- How to Categorise Roles | Function, Authority, Responsibility, Capability, Scope and Accountability
- How to Categorise Sensitivity | Public, Internal, Confidential, Restricted, Critical and Regulated
- How to Categorise Values | Intrinsic, Instrumental, Ethical, Social, Institutional and Non-Negotiable
