How to Classify New and Unknown Things | Novelty, Open Sets, Other and Taxonomy Expansion

The hardest classification problem is not choosing among known categories. It is recognising when none of the known categories is good enough.

A new technology arrives. A biological specimen does not fit familiar species patterns. A document type appears that an archive has never seen. A student gives an answer reflecting a misconception outside the existing error codes. A machine-learning system receives data from a new population. A hybrid organisation combines roles the taxonomy once kept separate.

A closed classifier will still choose the nearest familiar category. An open classification system can say: this may be new.

Quick answer: how should new or unknown things be classified?

  1. Check whether the item fits any existing category strongly enough.
  2. Compare the best candidate with alternatives.
  3. Detect low similarity, failed rules or out-of-distribution signals.
  4. Return unknown, novel or needs review instead of forcing a label.
  5. Preserve the evidence and candidate categories.
  6. Place genuinely unresolved items in a monitored holding state rather than a permanent junk bucket.
  7. Look for recurring patterns among novel cases.
  8. Create a new category only when the pattern is stable and useful.
  9. Version the taxonomy and crosswalk old records if the new category changes previous classifications.

This extends How to Categorise With Uncertainty from uncertain membership inside a known system to the deeper case where the system itself may be incomplete.


1. Closed-set classification assumes the answer exists

A closed-set classifier receives a fixed list of categories and must choose among them.

2. The nearest known category may still be wrong

If a new object is unlike every training example, “nearest” only means least distant, not truly appropriate.

3. Open-set recognition allows rejection

An open-set system can refuse known labels when fit is too weak.

4. Rejection is a valid classification outcome

“None of the known categories” can be more accurate than a confident wrong answer.

5. Unknown is not the same as Other

Unknown means the system lacks enough knowledge. Other means the item is understood well enough to know it sits outside named categories.

6. Novel is different again

A novel item contains a pattern sufficiently different from known cases that it may represent a new category, subtype or phenomenon.

7. Out-of-distribution means the input population changed

An item can be valid and familiar in the world while still falling outside the distribution used to build the classifier.

8. Novelty can be feature-based

The item may contain measurements, structures or behaviours far from known examples.

9. Novelty can be relational

An entity may participate in an unprecedented network role or relationship pattern.

10. Novelty can be semantic

New language or combinations of concepts may not map cleanly onto the existing vocabulary.

11. Novelty can be structural

A new object type may reveal that the ontology is missing an entire entity class rather than merely a leaf category.

12. Low confidence is one novelty signal

Repeated low confidence can indicate boundary ambiguity or absence of a suitable category.

13. Low similarity is another signal

If an item lies far from all prototypes or labelled examples, similarity-based systems should consider rejection.

14. Small margins are different from low fit

A case can fit two known categories nearly equally well. That is ambiguity, not necessarily novelty.

15. Rule failure can reveal novelty

If an item violates every known rule path yet remains clearly in scope, the taxonomy may be incomplete.

16. Impossible combinations can be meaningful

A supposedly impossible mix of features may expose bad data, a rare exception or a genuinely new phenomenon. Investigate before discarding it.

17. Create a review state for novelty

Use states such as novel candidate, out of distribution, taxonomy gap or needs domain review.

18. Do not hide novelty inside Miscellaneous

A permanent catch-all prevents the system from learning what it does not understand.

19. A holding area should be monitored

Track counts, themes, time and reviewer decisions for unresolved items.

20. One strange case does not justify a new category

Categories should represent recurring or operationally important distinctions, not every anomaly.

21. Repetition turns anomalies into signals

If similar novel cases recur, cluster them and ask whether they share a stable meaning.

22. Candidate categories need definitions

Before promoting a cluster into the taxonomy, write its purpose, admission criteria, exclusions and examples.

23. New categories need negative space

Identify the nearest existing category and explain why the new one deserves separation.

24. A subtype may be enough

Not every novel pattern requires a new top-level branch. It may fit as a child of an existing concept.

25. A new facet may be better than a new type

If novelty arises from a new independent dimension, add a facet rather than multiplying category combinations.

26. A new relationship may be the missing structure

Sometimes the object types are already correct but the ontology lacks a relationship needed to describe the new case.

27. New terminology may only require vocabulary expansion

A new word can be a synonym for an old concept rather than evidence for a new category.

28. Crosswalk candidate meaning before adding it

Search internal and external schemes to see whether the concept already exists under another name.

29. Novelty detection can be biased

Underrepresented groups may look novel simply because the reference data were narrow.

30. Representative data reduce false novelty

Expand reference examples before concluding that unfamiliarity equals a new category.

31. Novelty can be valuable information

Unexpected cases can reveal new science, new markets, new risks, new behaviours or hidden flaws in the existing taxonomy.

32. Novelty should feed governance

Recurring new cases should generate evidence-based taxonomy change requests rather than informal labels.

33. Taxonomy expansion needs versioning

When a new category is published, assign a new schema version and effective date.

34. Old records may need back-classification

Search historical unresolved or broad-category records to see whether some now fit the new concept.

35. Back-classification should preserve original history

Store the new classification alongside the old one with the reason and version change.

36. AI systems need explicit reject options

A model forced to choose among known labels cannot demonstrate open-set intelligence. Give it an allowed unknown or review outcome.

37. Human review should inspect clusters of rejected cases

Reviewing rejected items one by one solves immediate cases. Reviewing them as a population helps the taxonomy learn.

38. A practical novelty protocol

  1. Score fit to known categories.
  2. Measure margin among candidates.
  3. Check rule and ontology constraints.
  4. Flag weak-fit cases.
  5. Preserve candidate labels and evidence.
  6. Route to review.
  7. Cluster recurring novel cases.
  8. Test whether a new category, facet, relationship or synonym is needed.
  9. Approve through governance.
  10. Version and back-classify carefully.

39. The taxonomy should have an outside

A system that claims every possible future object already belongs somewhere cannot tell the difference between knowledge and forced familiarity.

40. The deeper idea

Classification is not only the art of naming what we know. It is also the discipline of recognising where our map ends.

The moment a system can say “I do not have a good category for this yet,” it becomes capable of learning a new category honestly.

Final answer

Classify new and unknown things by allowing rejection, novelty and out-of-distribution states. Do not force the nearest label. Preserve evidence, monitor unresolved cases, cluster recurring novelty, and expand the taxonomy only when a stable, useful distinction emerges. Version every expansion and preserve historical classifications.


Continue through the series

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.