Token Granularity | How Characters, Subwords, Words, Phrases and Concepts Create Different Model Worlds

Granularity is the scale at which a representation decides what counts as one unit. Change the scale and the model receives a different world.

Characters preserve fine detail but create long sequences. Whole words are intuitive but open-ended. Subwords trade vocabulary size against sequence length. Phrases compress larger patterns. Learned concepts can become even more abstract but risk hiding surface distinctions.

This article extends the eduKateSingapore Representation and Tokenisation series.

The Granularity Ladder

BYTE
→ CHARACTER
→ SUBWORD
→ WORD
→ PHRASE
→ SENTENCE
→ CHUNK
→ CONCEPT / EVENT / OBJECT

FINER UNITS = MORE DETAIL, LONGER SEQUENCES
COARSER UNITS = MORE COMPRESSION, STRONGER ASSUMPTIONS

1. Granularity Is a Representation Choice

The source does not announce one universally correct token scale. A sentence can be represented as bytes, characters, subwords, words, phrases or a single sentence embedding.

Each representation preserves different distinctions and supports different operations.

2. Fine Granularity Preserves Local Detail

Bytes and characters keep spelling, punctuation and unusual symbols close to the source. Rare words remain representable because they can always be decomposed.

The price is longer sequences and more steps required to build larger structure.

3. Coarse Granularity Compresses Repeated Structure

A whole phrase token can replace several words with one identity. A sentence embedding can represent hundreds of characters as one vector.

The price is that internal detail becomes less directly addressable.

4. Byte-Level Units Maximise Coverage

Bytes provide a tiny universal machine alphabet. Any valid encoded text can be represented as bytes, making unknown-character failure unlikely.

See Byte-Level Tokenisation.

5. Character-Level Units Align More Closely With Writing Systems

Characters are easier for humans to interpret than bytes and often preserve spelling structure cleanly.

But Unicode complicates the idea of “one character” because one displayed grapheme can contain multiple code points.

6. Character Models Face Long Sequences

A 1,000-word document can require several thousand character positions. Long sequences increase compute and make distant relationships harder to model efficiently.

Fine granularity spends compute to avoid vocabulary complexity.

7. Word-Level Units Are Human-Intuitive

Words feel like natural units because writing systems and dictionaries make them visible. Many classical NLP systems therefore worked directly at word level.

But natural language is productive, so whole-word vocabularies encounter names, inflections, compounds and new terms endlessly.

8. Word-Level Vocabularies Face the Open-Vocabulary Problem

A finite vocabulary cannot contain every possible word form. Unknown-word tokens collapse distinct strings and destroy reconstructability.

This is one reason subword tokenisation became dominant in modern language models.

9. Subwords Occupy the Middle Ground

Subwords keep common strings compact while composing rare strings from reusable parts. They reduce sequence length compared with characters and reduce vocabulary explosion compared with whole words.

BPE, WordPiece and Unigram implement this compromise differently.

10. Subword Granularity Shares Statistical Evidence

If educate, education and educational share subword pieces, training evidence can partly transfer across forms.

Yet statistical pieces do not always align perfectly with linguistic morphology.

11. Morphological Granularity Is a Different Objective

A linguistically motivated tokenizer might try to isolate roots, prefixes and suffixes explicitly. This can improve interpretability and compositionality in some languages.

But morphology itself can be ambiguous and language-specific.

12. Phrases Compress Multi-Word Regularities

Frequent expressions such as New York or machine learning can behave as coherent units in some tasks. Phrase-level tokens or learned spans reduce sequence length and preserve collocations.

The risk is over-compressing phrases whose internal words matter in other contexts.

13. Multiword Expressions Challenge Word-Level Independence

Kick the bucket cannot be interpreted purely by combining ordinary meanings of kick and bucket. Idioms show that useful semantic units can be larger than words.

Granularity should sometimes follow function rather than whitespace.

14. Named Entities Can Be Coarser Than Words

National University of Singapore contains several words but refers to one entity. Entity-aware systems can represent the full name with one canonical ID while preserving the original token sequence.

This creates two simultaneous granularities: lexical and entity-level.

15. Sentences Are Useful Semantic Units for Retrieval

A sentence can often express one proposition or compact reasoning step. Sentence embeddings therefore provide useful coarse representations for semantic search.

But a single sentence can contain several claims, while one claim can span several sentences.

16. Paragraphs Preserve Local Argument Structure

Paragraphs can provide more context than sentences and are often useful retrieval units. They preserve definitions, examples and qualifications together.

They are still editorial rather than universal semantic units.

17. Chunks Are Operational Units, Not Natural Units

Retrieval systems often divide documents into chunks of hundreds of tokens. The chunk is sized for indexing and context, not because the source naturally contains 500-token concepts.

See Token Budgets and Chunking.

18. Concepts Can Be Coarser Than Text Spans

The concept of photosynthesis can appear across many sentences, diagrams and equations. A concept-level representation can aggregate evidence across those forms.

Concept tokens therefore require stronger semantic assumptions than surface tokens.

19. Events Can Be Better Units Than Sentences

A news story may describe one event across several paragraphs. Event extraction can create a structured token-like object containing actors, time, place and action.

Higher granularity can make reasoning easier when the task concerns world events rather than wording.

20. Objects Can Be Better Units Than Image Patches

A vision model may process fixed patches, while another system aggregates patches into object representations. An object token can compress many pixels into one entity-level unit.

See Visual Tokenisation.

21. Audio Can Be Tokenised at Several Granularities Too

Waveform samples, frames, phoneme-like units, semantic speech tokens and codec tokens each preserve different temporal detail.

See Audio Tokenisation.

22. Fine Units Increase Sequence Length

If a word is split into four subwords instead of one token, the model uses four sequence positions. If an image uses 14×14 patches instead of 32×32 patches, the visual sequence grows sharply.

Granularity is directly connected to compute.

23. Coarse Units Increase Vocabulary or Encoder Complexity

Representing many phrases or concepts directly requires a larger vocabulary or a separate mechanism that discovers and encodes those units.

Sequence compression often moves complexity elsewhere in the system.

24. Granularity Changes the Learning Problem

A character model must learn word structure from long local sequences. A word model begins with word identities but must handle unknown forms. A concept model begins from abstract units but depends on a reliable concept extractor.

Different granularities relocate where intelligence must be learned.

25. Granularity Changes Error Types

Fine-grained systems can make spelling and boundary errors. Coarse-grained systems can make category and abstraction errors. A wrong character is local; a wrong concept assignment can distort an entire reasoning chain.

Compression increases the consequence of each unit.

26. Granularity Changes Interpretability

Human-readable word tokens are easy to inspect. Byte tokens are difficult to interpret. Concept tokens may be meaningful if explicitly labelled, but learned latent units can be opaque.

Operational convenience and human interpretability do not always align.

27. Dynamic Granularity Can Adapt to the Input

A system does not have to use one fixed scale everywhere. Common words can remain coarse while rare words decompose. Large document sections can be represented coarsely until a query requires finer detail.

Adaptive representation spends resolution where it is useful.

28. Hierarchical Representation Uses Several Scales at Once

A document can be represented as characters, tokens, sentences, paragraphs, sections and an overall document vector simultaneously.

Higher levels compress structure; lower levels preserve detail for verification.

29. Hierarchy Protects World Return

A model can reason at concept level and still return to the exact source sentence when evidence must be checked.

Strong systems preserve bridges between coarse abstraction and fine evidence.

30. Coarse Representations Need Provenance

A summary token such as “policy changed” is dangerous without a route back to which policy, when, where and according to which source.

As granularity becomes coarser, provenance becomes more important.

31. Multilingual Granularity Is Uneven

One language may be represented mostly through whole-word-like tokens while another uses smaller subwords under the same tokenizer.

Equal model context therefore provides unequal semantic granularity across languages.

32. Code Has Its Own Natural Granularity

Characters are too fine for most program reasoning, while entire files are too coarse. Identifiers, expressions, statements, functions and abstract-syntax-tree nodes provide useful intermediate levels.

Code systems often benefit from combining lexical and structural granularities.

33. Mathematics Also Has Multiple Granularities

An equation can be seen as characters, symbols, terms, subexpressions, identities or transformations. Students become stronger as they learn to chunk low-level symbols into higher-level mathematical objects.

Expertise is partly a change in representational granularity.

34. Teaching Moves Learners Between Scales

A teacher may zoom into one word, then a sentence, then a paragraph, then the argument. In mathematics, instruction can move from symbol to term to equation to model.

Learning depends on choosing the scale at which the next distinction becomes visible.

35. Evaluation Should Sweep Granularity

Instead of assuming one tokenizer scale is optimal, compare alternatives on sequence length, accuracy, interpretability, multilingual behaviour and downstream tasks.

The best granularity is empirical and receiver-specific.

36. The Granularity Audit

  1. What is the smallest unit preserved?
  2. What is the largest unit represented directly?
  3. How does unit size affect sequence length?
  4. How does it affect vocabulary or encoder size?
  5. What internal distinctions disappear at coarser scales?
  6. What long-range relationships become easier?
  7. Which unit scale aligns with the receiver’s task?
  8. Are multiple granularities available simultaneously?
  9. Can coarse representations return to fine source evidence?
  10. How does granularity vary across languages?
  11. What errors become more likely at each scale?
  12. Has downstream performance been measured across alternatives?

37. What Students Should Remember

38. The Deep Principle

Every representation has a zoom level. At fine scale we see detail without much abstraction. At coarse scale we see structure while local variation disappears.

Choose a token scale too small and the system drowns in pieces. Choose it too large and the system mistakes compression for understanding. Good granularity preserves the smallest distinctions the current job still needs.

Continue the Representation & Tokenisation Series

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading