Granularity is the scale at which a representation decides what counts as one unit. Change the scale and the model receives a different world.
Characters preserve fine detail but create long sequences. Whole words are intuitive but open-ended. Subwords trade vocabulary size against sequence length. Phrases compress larger patterns. Learned concepts can become even more abstract but risk hiding surface distinctions.
This article extends the eduKateSingapore Representation and Tokenisation series.
The Granularity Ladder
BYTE → CHARACTER → SUBWORD → WORD → PHRASE → SENTENCE → CHUNK → CONCEPT / EVENT / OBJECT FINER UNITS = MORE DETAIL, LONGER SEQUENCES COARSER UNITS = MORE COMPRESSION, STRONGER ASSUMPTIONS
1. Granularity Is a Representation Choice
The source does not announce one universally correct token scale. A sentence can be represented as bytes, characters, subwords, words, phrases or a single sentence embedding.
Each representation preserves different distinctions and supports different operations.
2. Fine Granularity Preserves Local Detail
Bytes and characters keep spelling, punctuation and unusual symbols close to the source. Rare words remain representable because they can always be decomposed.
The price is longer sequences and more steps required to build larger structure.
3. Coarse Granularity Compresses Repeated Structure
A whole phrase token can replace several words with one identity. A sentence embedding can represent hundreds of characters as one vector.
The price is that internal detail becomes less directly addressable.
4. Byte-Level Units Maximise Coverage
Bytes provide a tiny universal machine alphabet. Any valid encoded text can be represented as bytes, making unknown-character failure unlikely.
5. Character-Level Units Align More Closely With Writing Systems
Characters are easier for humans to interpret than bytes and often preserve spelling structure cleanly.
But Unicode complicates the idea of “one character” because one displayed grapheme can contain multiple code points.
6. Character Models Face Long Sequences
A 1,000-word document can require several thousand character positions. Long sequences increase compute and make distant relationships harder to model efficiently.
Fine granularity spends compute to avoid vocabulary complexity.
7. Word-Level Units Are Human-Intuitive
Words feel like natural units because writing systems and dictionaries make them visible. Many classical NLP systems therefore worked directly at word level.
But natural language is productive, so whole-word vocabularies encounter names, inflections, compounds and new terms endlessly.
8. Word-Level Vocabularies Face the Open-Vocabulary Problem
A finite vocabulary cannot contain every possible word form. Unknown-word tokens collapse distinct strings and destroy reconstructability.
This is one reason subword tokenisation became dominant in modern language models.
9. Subwords Occupy the Middle Ground
Subwords keep common strings compact while composing rare strings from reusable parts. They reduce sequence length compared with characters and reduce vocabulary explosion compared with whole words.
BPE, WordPiece and Unigram implement this compromise differently.
10. Subword Granularity Shares Statistical Evidence
If educate, education and educational share subword pieces, training evidence can partly transfer across forms.
Yet statistical pieces do not always align perfectly with linguistic morphology.
11. Morphological Granularity Is a Different Objective
A linguistically motivated tokenizer might try to isolate roots, prefixes and suffixes explicitly. This can improve interpretability and compositionality in some languages.
But morphology itself can be ambiguous and language-specific.
12. Phrases Compress Multi-Word Regularities
Frequent expressions such as New York or machine learning can behave as coherent units in some tasks. Phrase-level tokens or learned spans reduce sequence length and preserve collocations.
The risk is over-compressing phrases whose internal words matter in other contexts.
13. Multiword Expressions Challenge Word-Level Independence
Kick the bucket cannot be interpreted purely by combining ordinary meanings of kick and bucket. Idioms show that useful semantic units can be larger than words.
Granularity should sometimes follow function rather than whitespace.
14. Named Entities Can Be Coarser Than Words
National University of Singapore contains several words but refers to one entity. Entity-aware systems can represent the full name with one canonical ID while preserving the original token sequence.
This creates two simultaneous granularities: lexical and entity-level.
15. Sentences Are Useful Semantic Units for Retrieval
A sentence can often express one proposition or compact reasoning step. Sentence embeddings therefore provide useful coarse representations for semantic search.
But a single sentence can contain several claims, while one claim can span several sentences.
16. Paragraphs Preserve Local Argument Structure
Paragraphs can provide more context than sentences and are often useful retrieval units. They preserve definitions, examples and qualifications together.
They are still editorial rather than universal semantic units.
17. Chunks Are Operational Units, Not Natural Units
Retrieval systems often divide documents into chunks of hundreds of tokens. The chunk is sized for indexing and context, not because the source naturally contains 500-token concepts.
See Token Budgets and Chunking.
18. Concepts Can Be Coarser Than Text Spans
The concept of photosynthesis can appear across many sentences, diagrams and equations. A concept-level representation can aggregate evidence across those forms.
Concept tokens therefore require stronger semantic assumptions than surface tokens.
19. Events Can Be Better Units Than Sentences
A news story may describe one event across several paragraphs. Event extraction can create a structured token-like object containing actors, time, place and action.
Higher granularity can make reasoning easier when the task concerns world events rather than wording.
20. Objects Can Be Better Units Than Image Patches
A vision model may process fixed patches, while another system aggregates patches into object representations. An object token can compress many pixels into one entity-level unit.
See Visual Tokenisation.
21. Audio Can Be Tokenised at Several Granularities Too
Waveform samples, frames, phoneme-like units, semantic speech tokens and codec tokens each preserve different temporal detail.
See Audio Tokenisation.
22. Fine Units Increase Sequence Length
If a word is split into four subwords instead of one token, the model uses four sequence positions. If an image uses 14×14 patches instead of 32×32 patches, the visual sequence grows sharply.
Granularity is directly connected to compute.
23. Coarse Units Increase Vocabulary or Encoder Complexity
Representing many phrases or concepts directly requires a larger vocabulary or a separate mechanism that discovers and encodes those units.
Sequence compression often moves complexity elsewhere in the system.
24. Granularity Changes the Learning Problem
A character model must learn word structure from long local sequences. A word model begins with word identities but must handle unknown forms. A concept model begins from abstract units but depends on a reliable concept extractor.
Different granularities relocate where intelligence must be learned.
25. Granularity Changes Error Types
Fine-grained systems can make spelling and boundary errors. Coarse-grained systems can make category and abstraction errors. A wrong character is local; a wrong concept assignment can distort an entire reasoning chain.
Compression increases the consequence of each unit.
26. Granularity Changes Interpretability
Human-readable word tokens are easy to inspect. Byte tokens are difficult to interpret. Concept tokens may be meaningful if explicitly labelled, but learned latent units can be opaque.
Operational convenience and human interpretability do not always align.
27. Dynamic Granularity Can Adapt to the Input
A system does not have to use one fixed scale everywhere. Common words can remain coarse while rare words decompose. Large document sections can be represented coarsely until a query requires finer detail.
Adaptive representation spends resolution where it is useful.
28. Hierarchical Representation Uses Several Scales at Once
A document can be represented as characters, tokens, sentences, paragraphs, sections and an overall document vector simultaneously.
Higher levels compress structure; lower levels preserve detail for verification.
29. Hierarchy Protects World Return
A model can reason at concept level and still return to the exact source sentence when evidence must be checked.
Strong systems preserve bridges between coarse abstraction and fine evidence.
30. Coarse Representations Need Provenance
A summary token such as “policy changed” is dangerous without a route back to which policy, when, where and according to which source.
As granularity becomes coarser, provenance becomes more important.
31. Multilingual Granularity Is Uneven
One language may be represented mostly through whole-word-like tokens while another uses smaller subwords under the same tokenizer.
Equal model context therefore provides unequal semantic granularity across languages.
32. Code Has Its Own Natural Granularity
Characters are too fine for most program reasoning, while entire files are too coarse. Identifiers, expressions, statements, functions and abstract-syntax-tree nodes provide useful intermediate levels.
Code systems often benefit from combining lexical and structural granularities.
33. Mathematics Also Has Multiple Granularities
An equation can be seen as characters, symbols, terms, subexpressions, identities or transformations. Students become stronger as they learn to chunk low-level symbols into higher-level mathematical objects.
Expertise is partly a change in representational granularity.
34. Teaching Moves Learners Between Scales
A teacher may zoom into one word, then a sentence, then a paragraph, then the argument. In mathematics, instruction can move from symbol to term to equation to model.
Learning depends on choosing the scale at which the next distinction becomes visible.
35. Evaluation Should Sweep Granularity
Instead of assuming one tokenizer scale is optimal, compare alternatives on sequence length, accuracy, interpretability, multilingual behaviour and downstream tasks.
The best granularity is empirical and receiver-specific.
36. The Granularity Audit
- What is the smallest unit preserved?
- What is the largest unit represented directly?
- How does unit size affect sequence length?
- How does it affect vocabulary or encoder size?
- What internal distinctions disappear at coarser scales?
- What long-range relationships become easier?
- Which unit scale aligns with the receiver’s task?
- Are multiple granularities available simultaneously?
- Can coarse representations return to fine source evidence?
- How does granularity vary across languages?
- What errors become more likely at each scale?
- Has downstream performance been measured across alternatives?
37. What Students Should Remember
- Granularity is the scale of representation.
- Finer units preserve detail but create longer sequences.
- Coarser units compress structure but make stronger assumptions.
- Subwords balance words and characters.
- Phrases, entities, events and concepts can be higher-level tokens.
- No one granularity is best for every task.
- Hierarchical systems can preserve several scales at once.
38. The Deep Principle
Every representation has a zoom level. At fine scale we see detail without much abstraction. At coarse scale we see structure while local variation disappears.
Choose a token scale too small and the system drowns in pieces. Choose it too large and the system mistakes compression for understanding. Good granularity preserves the smallest distinctions the current job still needs.
