A category is only as strong as the criteria that decide who gets in and who stays out.
Labels are easy to invent. Criteria are harder. “Large company”, “advanced student”, “high risk”, “urban area”, “historical document”, “urgent case” and “reliable source” all sound understandable until two careful people apply them to the same borderline example and disagree.
Good classification criteria convert a vague distinction into an inspectable rule without pretending every domain has a perfectly sharp boundary.
Quick answer: how do you choose classification criteria?
- Start with the purpose of the classification.
- Define the object being classified.
- Identify features that genuinely discriminate among categories.
- Prefer direct evidence over convenient proxies where possible.
- Separate necessary, sufficient and merely typical features.
- Set thresholds only when the job requires them.
- Test just-inside and just-outside boundary cases.
- Check reliability across independent classifiers.
- Measure error costs, not only agreement.
- Review bias and subgroup performance.
- Version criteria when evidence or purpose changes.
This article deepens the rule-writing stage in How to Categorise Anything and connects directly to How to Test a Classification System.
1. Criteria begin with purpose
The same object may need different criteria for storage, teaching, regulation, retrieval or diagnosis.
A criterion is useful only relative to a job.
2. Define the unit before the feature
Are you classifying a person, event, document, organisation, process, claim or physical object?
The relevant criteria depend on the entity type.
3. A useful criterion separates categories
If a feature is shared equally by every candidate category, it cannot do much discriminating work.
Good criteria identify meaningful differences.
4. Distinguish defining from typical features
Birds often fly, but flight does not define birds. Whales live in water, but aquatic life does not define mammals.
Typical features help recognition. Defining features establish membership.
5. Necessary criteria must always hold
A necessary condition is one that every valid member must satisfy.
If a supposed member can violate the condition and still belong, the condition was not actually necessary.
6. Sufficient criteria establish membership
A sufficient set of conditions is enough to place an item in a category.
Some domains support clean sufficient rules; many everyday categories do not.
7. Typicality is useful when strict rules fail
For fuzzy categories, criteria may describe a prototype rather than a hard logical boundary.
Record that the category is similarity-based rather than pretending the rule is exact.
8. Direct evidence is usually stronger
If the target concept can be measured directly, use that evidence instead of an indirect stand-in.
Direct evidence reduces proxy error.
9. Proxies need explicit justification
Sometimes the desired feature is too expensive, slow or impossible to observe directly.
A proxy can be useful, but document what it approximates and where it is known to fail.
10. Observable criteria improve reproducibility
“Looks important” is difficult to audit. “Appears in three official records” is easier to reproduce.
Operational criteria should point toward evidence another reviewer can inspect.
11. Criteria can be qualitative
Not every useful distinction has a numerical measurement.
Genre, legal reasoning, archival significance and many cultural categories require structured qualitative judgement.
12. Qualitative does not mean arbitrary
Qualitative criteria can still use definitions, examples, counterexamples, evidence requirements and review processes.
13. Thresholds convert continua into categories
Height, age, income, risk score and temperature are continuous or ordered quantities.
A threshold creates a categorical boundary for an operational reason.
14. Thresholds are conventions unless nature provides a discontinuity
Two cases immediately on opposite sides of a threshold may be nearly identical.
The category may still be administratively useful, but the convention should be acknowledged.
15. Test the exact threshold
For every numerical boundary, test below, exactly at and above the threshold.
This reveals ambiguity in inclusive and exclusive wording.
16. Avoid hidden compound criteria
A label like “high-value customer” may secretly combine spending, frequency, tenure and profitability.
Write the components explicitly so the category can be audited.
17. Decide whether all criteria are required
Some rules use AND logic: A and B and C must all hold. Others use OR logic or weighted evidence.
The combination rule is part of the category definition.
18. Weighted criteria create scoring systems
When several weak signals contribute to a classification, a score may be more appropriate than one decisive feature.
Weights should be validated rather than chosen for cosmetic precision.
19. Scoring introduces another threshold
A score still needs a rule for when membership begins.
Test threshold sensitivity and error trade-offs.
20. Criteria should minimise circularity
“A strategic project is one considered strategic” cannot be independently tested.
A definition should add discriminating information rather than repeat the label.
21. Criteria need exclusion rules too
It is often easier to understand a boundary by knowing what looks similar but does not belong.
Write explicit exclusions for common near misses.
22. Counterexamples stress-test the rule
A strong criterion survives awkward examples or explains clearly why the awkward case is an exception.
See the site’s science guide Understanding Why Classification Needs Clear Criteria for the educational version of this principle.
23. Criteria need evidence rules
What evidence is enough to establish the criterion?
One self-report? An official record? Two independent sources? A measurement? Expert inspection?
24. Source quality matters
Two pieces of evidence do not necessarily carry equal weight.
Classification should distinguish weak from authoritative sources where the stakes justify it.
25. Missing evidence should not force failure
If a criterion cannot be evaluated because evidence is absent, return unknown or provisional where appropriate rather than treating absence as false.
See How to Categorise With Uncertainty.
26. Criteria should be understandable to classifiers
A theoretically precise rule may still fail if operational users cannot apply it consistently.
Use plain explanations and examples alongside technical definitions.
27. Measure inter-rater agreement
Give the same cases to independent classifiers.
Persistent disagreement shows where the criteria need clarification or the domain is genuinely ambiguous.
28. Agreement does not prove validity
Everyone can apply a poor rule consistently.
Also test whether the criterion predicts, retrieves or routes what the system actually cares about.
29. Measure false inclusion
Which non-members are admitted by the criterion?
False inclusions reveal overly broad rules.
30. Measure false exclusion
Which valid members are rejected?
False exclusions reveal overly narrow rules or missing evidence pathways.
31. Error costs may be asymmetric
Missing a hazardous item may cost more than temporarily over-classifying a safe one.
Choose criteria and thresholds with consequence in mind.
32. Criteria can create bias
A convenient feature may work differently across groups or contexts.
Audit proxy quality, measurement availability and error patterns using How Classification Bias Works.
33. Criteria can drift
The same feature may lose discriminating power as technology or behaviour changes.
Review criteria against recent data rather than assuming historical usefulness remains permanent.
34. Criteria should be versioned
If the admission rule changes, the same label may mean something different under the new version.
Store effective dates and crosswalk implications.
35. AI needs explicit criteria too
A language model may infer category meaning from names, but explicit definitions and examples make automated classification more stable.
Do not rely on the label string as the whole specification.
36. Evidence spans improve machine auditability
Where possible, require automated classifiers to point to the features or text that support the assignment.
This helps reviewers distinguish genuine reasoning from shortcut patterns.
37. A practical category-criteria card
- category ID;
- purpose;
- definition;
- necessary criteria;
- sufficient criteria;
- typical features;
- exclusion criteria;
- thresholds;
- evidence requirements;
- known proxies;
- examples;
- counterexamples;
- uncertainty rule;
- version.
38. A complete criteria test
- Test obvious positives.
- Test obvious negatives.
- Test one-condition failures.
- Test boundary values.
- Test counterexamples.
- Test missing evidence.
- Test independent agreement.
- Measure false inclusion and exclusion.
- Inspect subgroup errors.
- Test downstream usefulness.
39. The best criterion is not always the most complex
A simpler criterion that is observable, reliable and strongly discriminating may outperform a complex score that nobody can explain or maintain.
40. The deeper idea
Criteria are where category philosophy becomes operational behaviour.
A category name tells you what the box is called. Criteria tell you why this object belongs inside it.
Final answer
Choose classification criteria that serve the task, discriminate meaningfully, rely on appropriate evidence and survive boundary testing. Separate defining from typical features. Justify proxies and thresholds. Measure reliability, false inclusion, false exclusion and bias. Preserve uncertainty and version criteria when the domain changes.
