The hardest classification problem is not choosing among known categories. It is recognising when none of the known categories is good enough.
A new technology arrives. A biological specimen does not fit familiar species patterns. A document type appears that an archive has never seen. A student gives an answer reflecting a misconception outside the existing error codes. A machine-learning system receives data from a new population. A hybrid organisation combines roles the taxonomy once kept separate.
A closed classifier will still choose the nearest familiar category. An open classification system can say: this may be new.
Quick answer: how should new or unknown things be classified?
- Check whether the item fits any existing category strongly enough.
- Compare the best candidate with alternatives.
- Detect low similarity, failed rules or out-of-distribution signals.
- Return unknown, novel or needs review instead of forcing a label.
- Preserve the evidence and candidate categories.
- Place genuinely unresolved items in a monitored holding state rather than a permanent junk bucket.
- Look for recurring patterns among novel cases.
- Create a new category only when the pattern is stable and useful.
- Version the taxonomy and crosswalk old records if the new category changes previous classifications.
This extends How to Categorise With Uncertainty from uncertain membership inside a known system to the deeper case where the system itself may be incomplete.
1. Closed-set classification assumes the answer exists
A closed-set classifier receives a fixed list of categories and must choose among them.
2. The nearest known category may still be wrong
If a new object is unlike every training example, “nearest” only means least distant, not truly appropriate.
3. Open-set recognition allows rejection
An open-set system can refuse known labels when fit is too weak.
4. Rejection is a valid classification outcome
“None of the known categories” can be more accurate than a confident wrong answer.
5. Unknown is not the same as Other
Unknown means the system lacks enough knowledge. Other means the item is understood well enough to know it sits outside named categories.
6. Novel is different again
A novel item contains a pattern sufficiently different from known cases that it may represent a new category, subtype or phenomenon.
7. Out-of-distribution means the input population changed
An item can be valid and familiar in the world while still falling outside the distribution used to build the classifier.
8. Novelty can be feature-based
The item may contain measurements, structures or behaviours far from known examples.
9. Novelty can be relational
An entity may participate in an unprecedented network role or relationship pattern.
10. Novelty can be semantic
New language or combinations of concepts may not map cleanly onto the existing vocabulary.
11. Novelty can be structural
A new object type may reveal that the ontology is missing an entire entity class rather than merely a leaf category.
12. Low confidence is one novelty signal
Repeated low confidence can indicate boundary ambiguity or absence of a suitable category.
13. Low similarity is another signal
If an item lies far from all prototypes or labelled examples, similarity-based systems should consider rejection.
14. Small margins are different from low fit
A case can fit two known categories nearly equally well. That is ambiguity, not necessarily novelty.
15. Rule failure can reveal novelty
If an item violates every known rule path yet remains clearly in scope, the taxonomy may be incomplete.
16. Impossible combinations can be meaningful
A supposedly impossible mix of features may expose bad data, a rare exception or a genuinely new phenomenon. Investigate before discarding it.
17. Create a review state for novelty
Use states such as novel candidate, out of distribution, taxonomy gap or needs domain review.
18. Do not hide novelty inside Miscellaneous
A permanent catch-all prevents the system from learning what it does not understand.
19. A holding area should be monitored
Track counts, themes, time and reviewer decisions for unresolved items.
20. One strange case does not justify a new category
Categories should represent recurring or operationally important distinctions, not every anomaly.
21. Repetition turns anomalies into signals
If similar novel cases recur, cluster them and ask whether they share a stable meaning.
22. Candidate categories need definitions
Before promoting a cluster into the taxonomy, write its purpose, admission criteria, exclusions and examples.
23. New categories need negative space
Identify the nearest existing category and explain why the new one deserves separation.
24. A subtype may be enough
Not every novel pattern requires a new top-level branch. It may fit as a child of an existing concept.
25. A new facet may be better than a new type
If novelty arises from a new independent dimension, add a facet rather than multiplying category combinations.
26. A new relationship may be the missing structure
Sometimes the object types are already correct but the ontology lacks a relationship needed to describe the new case.
27. New terminology may only require vocabulary expansion
A new word can be a synonym for an old concept rather than evidence for a new category.
28. Crosswalk candidate meaning before adding it
Search internal and external schemes to see whether the concept already exists under another name.
29. Novelty detection can be biased
Underrepresented groups may look novel simply because the reference data were narrow.
30. Representative data reduce false novelty
Expand reference examples before concluding that unfamiliarity equals a new category.
31. Novelty can be valuable information
Unexpected cases can reveal new science, new markets, new risks, new behaviours or hidden flaws in the existing taxonomy.
32. Novelty should feed governance
Recurring new cases should generate evidence-based taxonomy change requests rather than informal labels.
33. Taxonomy expansion needs versioning
When a new category is published, assign a new schema version and effective date.
34. Old records may need back-classification
Search historical unresolved or broad-category records to see whether some now fit the new concept.
35. Back-classification should preserve original history
Store the new classification alongside the old one with the reason and version change.
36. AI systems need explicit reject options
A model forced to choose among known labels cannot demonstrate open-set intelligence. Give it an allowed unknown or review outcome.
37. Human review should inspect clusters of rejected cases
Reviewing rejected items one by one solves immediate cases. Reviewing them as a population helps the taxonomy learn.
38. A practical novelty protocol
- Score fit to known categories.
- Measure margin among candidates.
- Check rule and ontology constraints.
- Flag weak-fit cases.
- Preserve candidate labels and evidence.
- Route to review.
- Cluster recurring novel cases.
- Test whether a new category, facet, relationship or synonym is needed.
- Approve through governance.
- Version and back-classify carefully.
39. The taxonomy should have an outside
A system that claims every possible future object already belongs somewhere cannot tell the difference between knowledge and forced familiarity.
40. The deeper idea
Classification is not only the art of naming what we know. It is also the discipline of recognising where our map ends.
The moment a system can say “I do not have a good category for this yet,” it becomes capable of learning a new category honestly.
Final answer
Classify new and unknown things by allowing rejection, novelty and out-of-distribution states. Do not force the nearest label. Preserve evidence, monitor unresolved cases, cluster recurring novelty, and expand the taxonomy only when a stable, useful distinction emerges. Version every expansion and preserve historical classifications.
Continue through the series
- How to Categorise Anything | A General Framework for Classification
- How to Categorise With Uncertainty | Unknown, Probable, Disputed and Provisional
- How Dynamic Classification Systems Work | Change, Drift, Versioning and Reclassification
- How AI Classification Works | From Signals and Labels to Confidence, Review and Retrieval
