How Tokenisation Works | Boundaries, Vocabularies, Subwords and Reconstruction

Tokenisation is the boundary-making stage that turns an input representation into a sequence of units a computational system can address. It sounds like a small preprocessing step. In practice, it determines vocabulary size, sequence length, coverage, reconstruction behaviour, multilingual efficiency and part of the computational geometry that every later layer inherits.

This article sits underneath Representation and Tokenisation | How Information Becomes Units a System Can Work With and under the canonical World Representation & Cognitive Tools owner. The job here is narrower: follow the mechanism from source string to workable units and back again.

The Pipeline in One Line

SOURCE TEXT → ENCODING → NORMALISATION → PRE-TOKENISATION → SEGMENTATION → TOKEN VOCABULARY LOOKUP → TOKEN IDs → MODEL INPUT → OUTPUT IDs → DECODING → RECONSTRUCTED TEXT

Different tokenisers rearrange, combine or omit some stages, but this sequence is a useful control model. It lets us ask where a behaviour came from instead of blaming everything on the language model itself.

1. Begin With the Representation You Actually Have

A tokenizer does not receive meaning directly. It receives a representation: usually an encoded text string, sometimes after another layer has already extracted text from speech, images, documents or structured data. If optical character recognition mistakes a letter, the tokenizer faithfully tokenises the mistake. If a PDF extractor loses column order, the tokenizer receives the wrong sequence. If a transcription system drops a speaker label, tokenisation cannot restore that lost identity by itself.

Tokenisation must therefore be analysed as one stage in a longer chain. A clean tokenizer cannot repair an unclean source representation unless some separate correction mechanism exists.

2. Encoding: Characters Must Become Machine-Representable

Modern text systems typically work with Unicode, which assigns code points to characters and symbols across writing systems. But the same visible character sequence can sometimes have more than one underlying code-point sequence. This is why text that looks identical to a reader may fail a byte-level or code-point-level equality test.

Unicode normalisation addresses classes of equivalent representations. The Unicode Consortium specifies standard normalisation forms in Unicode Standard Annex #15. The lesson for tokenisation is not that every system should normalise everything identically. It is that the representation entering the tokenizer is already a technical choice.

3. Normalisation: Decide Which Differences Should Survive

Normalisation can lowercase text, standardise equivalent Unicode sequences, canonicalise whitespace or perform other transformations. Each operation changes what distinctions remain available. Lowercasing can improve vocabulary sharing across sentence positions, but it also removes a distinction that may matter for proper names, acronyms or case-sensitive code. Whitespace cleanup can simplify text, but it may destroy indentation that carries meaning in programming languages or formatted documents.

The correct question is not whether normalisation is good. It is: which distinctions are safe to collapse for this receiver and task?

4. Pre-Tokenisation: Establish Candidate Regions

Some tokenizer pipelines first split or mark the input around spaces, punctuation or other patterns before applying a learned subword model. Others treat raw text more directly. Pre-tokenisation can influence every downstream boundary because it may prohibit a learned token from crossing particular separators.

This is one reason tokenisers that appear to use the same broad algorithm can still produce different segmentations. The complete pipeline matters, not only the label BPE, WordPiece or Unigram.

5. The Vocabulary Is a Finite Set of Available Pieces

A tokenizer vocabulary is a controlled inventory of pieces the system knows how to identify. It can contain whole words, common fragments, punctuation-associated forms, bytes, characters and special control symbols. A larger vocabulary can represent frequent patterns compactly but increases the number of available token identities. A smaller vocabulary is easier to cover but tends to produce longer sequences.

LARGER VOCABULARY → MORE COMPACT COMMON PATTERNS → MORE TOKEN TYPES
SMALLER VOCABULARY → BROADER COMPOSITION FROM SMALL PIECES → LONGER SEQUENCES

6. Why Whole-Word Tokenisation Breaks at Scale

Whole-word tokenisation is attractive because human readers already recognise words. But language is productive. New names, compounds, inflections, spelling variants and specialist terms appear continuously. A word vocabulary large enough to include everything becomes impractical and still cannot guarantee coverage of tomorrow’s vocabulary.

The alternative is compositionality: represent uncommon words through smaller reusable parts. Subword tokenisation became a standard solution because it combines useful compression with broad coverage.

7. Byte Pair Encoding: Build Frequent Larger Pieces

Byte Pair Encoding, adapted for subword tokenisation, starts from small units and repeatedly combines frequent neighbouring pairs according to a learned merge process. Common strings can become single vocabulary items while rare strings remain decomposable into smaller units.

Rico Sennrich, Barry Haddow and Alexandra Birch’s 2016 paper Neural Machine Translation of Rare Words with Subword Units helped establish the value of this approach for open-vocabulary neural language processing. The key systems insight is that frequency can be used to choose which boundaries deserve compression.

BPE-style segmentation is useful, but it does not claim that every learned piece is a linguistic morpheme. A frequent character sequence may be computationally useful even when it has no clean grammatical status.

8. WordPiece: Another Route to Reusable Subwords

WordPiece is another influential subword family associated with BERT-style systems. It also builds a subword vocabulary, but its training and segmentation logic differs from standard BPE. Practical explanations are available in the Hugging Face tokenisation documentation.

The important educational point is that two subword tokenisers can receive the same visible text and choose different boundaries because their vocabularies and segmentation rules differ.

9. Unigram: Optimise a Probabilistic Vocabulary

Unigram tokenisation takes a probabilistic approach. It can begin with a large candidate inventory and estimate which pieces contribute most to representing the training data. Segmentation is selected under a learned probability model. This makes the underlying principle visible: tokenisation is optimisation under a vocabulary and objective, not discovery of one inevitable cut.

10. SentencePiece: Raw Text as the Starting Point

Taku Kudo and John Richardson described SentencePiece as a language-independent subword tokenizer and detokenizer that can train directly from raw sentences. This matters because many older pipelines assumed that whitespace-delimited words already existed as a preprocessing layer, an assumption that does not transfer cleanly across all writing systems.

11. Byte-Level Routes: Coverage From a Lower Layer

Byte-level approaches can provide broad coverage because encoded text can ultimately be expressed through bytes. This reduces dependence on a fixed character inventory and gives a route for unusual strings, mixed scripts and arbitrary text. The trade-off is that the pieces may be less intuitive from a human linguistic perspective.

12. Special Tokens Are Part of the Protocol

Many model pipelines reserve special tokens or identifiers for sequence boundaries, padding, masks, roles, separators or control states. These units may not correspond to visible source text at all. They exist because the tokenizer-model pair is also a communication protocol.

13. Tokens Become IDs

Once pieces are chosen, each token maps to an identifier. The identifier is usually an integer indexing the vocabulary. If token 4172 corresponds to a particular piece, the number 4172 does not contain the meaning of that piece. It is an address.

This matters because adjacent token IDs do not normally imply semantic similarity. Semantic structure is learned in later numerical representations, not in the arbitrary ordering of vocabulary IDs.

14. IDs Become Numerical Representations

The model maps token identities into numerical vectors or other internal representations suitable for learned computation. Position or ordering information is also incorporated because the sequence dog bites man is not equivalent to man bites dog.

15. Context Changes the Working Meaning

A token’s initial identity does not settle its final interpretation. Contextual sequence models repeatedly transform internal states using information from surrounding positions. The token representing cell behaves differently in discussions of biology, prisons, spreadsheets and batteries because the sequence around it alters the working representation.

16. Decoding Is the Return Path

A tokenizer usually supports decoding: converting token identities back into human-visible text. A well-designed tokenizer can make this route deterministic for ordinary inputs, but special tokens, normalisation choices or invalid sequences complicate the idea of perfect reconstruction.

17. Reversibility Is Not Semantic Fidelity

A tokenizer can reconstruct every character and still choose boundaries that are inconvenient for a downstream task. Reversibility answers whether the source can return. Semantic utility answers whether the chosen pieces support useful computation. These are different tests.

18. Whitespace Is Not Empty

Human readers often treat spaces as absence. Tokenizers may treat them as meaningful boundary signals or incorporate them into token pieces. In code, indentation can be structural. In poetry, spacing can carry form. In tables and fixed-width formats, whitespace can represent columns. Casual cleanup can therefore destroy information before tokenisation begins.

19. Punctuation Is Not Decorative

Punctuation can mark sentence structure, quotation, mathematical grouping, paths, URLs, programming syntax or emotive tone. A tokenizer may isolate punctuation, attach it to neighbouring text or learn multi-character combinations. In ordinary prose a comma may be easy to infer; in a regular expression or programming language, one mark can change the entire object.

20. Numbers Reveal the Difference Between Symbol and Concept

Humans learn that 2048 is one number with place-value structure. A tokenizer may represent the digit string using one piece or several. The model must therefore learn mathematical relationships through a representation that may not align with the conceptual decomposition taught in mathematics.

21. Code Has Different Natural Boundaries

Programming languages contain identifiers, operators, indentation, delimiters, literals and punctuation-dense expressions. A tokenizer built mainly from ordinary prose may fragment code differently from one trained with substantial code. The same principle applies to mathematics, chemistry, law and medicine: domain frequency changes which sequences become economical units.

22. Proper Names and Rare Terms Expose Coverage

A rare surname, place name or scientific term may be divided into several familiar pieces. This is not automatically failure. Subword systems are designed to represent unseen forms compositionally. But longer fragmentations can make exact copying, spelling-sensitive generation or entity matching more demanding.

23. Multilingual Tokenisation Is an Allocation Problem

A finite vocabulary allocates capacity across writing systems and patterns seen during tokenizer training. Languages with abundant representation may gain many compact frequent pieces. Others may require longer decompositions for similar amounts of human-visible text. Equal token budgets therefore do not always mean equal amounts of linguistic content.

24. Context Windows Are Measured After Tokenisation

Model context capacity is defined in model-side units, not human words or pages. A densely tokenised language, unusual codebase or punctuation-heavy dataset can consume context faster than ordinary prose. Capacity planning should therefore count with the actual tokenizer rather than rely on a fixed words-per-token rule.

25. Retrieval Adds Another Boundary Layer

Long-document retrieval often divides sources into chunks before embedding, indexing or model ingestion. Chunking is a second boundary system layered above tokenizer boundaries. A bad retrieval cut can separate a definition from its exception, a table heading from its rows or a condition from its qualifier. Tokenisation belongs to a wider family of segmentation problems.

26. Token Counts Belong to a Tokenizer-Text Pair

There is no intrinsic token count attached to a sentence independent of a tokenizer. Change the tokenizer and the count can change. Change Unicode representation, whitespace or punctuation and the segmentation may change again.

A token count is not a property of language alone. It is a property of an input representation processed by a particular tokenizer.

27. Training-Time and Inference-Time Compatibility Matters

A model learns on sequences produced by a particular tokenisation scheme. Using an incompatible vocabulary at inference can break the mapping between token IDs and learned parameters. The tokenizer and model are therefore not freely interchangeable accessories. They are part of one trained interface contract.

28. Version the Tokenizer

When reproducibility matters, record the tokenizer family, vocabulary and exact version or artifact. If a later system silently changes normalisation, special-token handling or vocabulary, historical token counts and model inputs may no longer be identical. Tokenizers belong in production version control just as other interfaces do.

29. Preserve Offsets Back to the Source

Many tokenizer APIs return offsets mapping tokens to character spans in source text. These are valuable for highlighting, annotation, entity tasks and debugging. Offsets make a general systems principle explicit: transformations are safer when they preserve a traceable route back to the source.

30. Test the Round Trip

A basic QA test is encode → decode → compare. The correct comparison depends on normalisation policy. If equivalent forms are intentionally collapsed, exact byte equality may not be the right invariant. The test must first declare what is required to survive.

31. Test Adversarial Boundaries

Do not test only ordinary prose. Test long numbers, mixed scripts, emoji sequences, uncommon names, repeated punctuation, zero-width characters, unusual whitespace, code, URLs, equations and malformed input. These cases reveal assumptions hidden by familiar sentences.

32. Do Not Infer Linguistic Truth From Token Boundaries

If a tokenizer splits a word into three pieces, that does not prove the language contains three morphemes there. If it keeps a phrase fragment together, that does not prove the fragment is one semantic unit. The tokenizer optimises a computational objective under a particular corpus and vocabulary budget.

33. Do Not Infer Understanding From Compact Tokens

A common word may be represented by one token while a rare technical term uses several. That says something about representation efficiency, not necessarily conceptual competence. Tokenisation is one constraint in the system, not a complete explanation of behaviour.

34. Human Expertise Also Changes Effective Tokens

Experts chunk information. A chess expert sees a formation where a novice sees isolated pieces. A mathematician sees a known structure where a beginner sees symbols. A fluent reader sees phrases where an early reader decodes letters. Good units reduce working-memory load while preserving structure needed for reasoning.

35. Tokenisation Debugging Checklist

  1. What exact source representation enters the tokenizer?
  2. What encoding and normalisation rules apply?
  3. Is there pre-tokenisation?
  4. Which vocabulary and tokenizer version are used?
  5. Which algorithm family performs segmentation?
  6. How are spaces, punctuation and special symbols handled?
  7. How are rare sequences represented?
  8. Can tokens map back to source offsets?
  9. Can encoding and decoding preserve the required invariants?
  10. How many tokens do representative languages and domains consume?
  11. What happens to long numbers, code, URLs and unusual scripts?
  12. Is the tokenizer exactly compatible with the model?

36. The Deeper Principle

Tokenisation is not merely text chopping. It is a negotiated boundary between human-visible representation and model-ready computation. The tokenizer decides which recurring forms deserve compact identities, which rare forms must be composed from smaller pieces and which differences are preserved or collapsed before the model begins its main work.

Before intelligence can operate on a world, the world must be represented. Before a representation can be computed over, it must be given workable structure. Every boundary that makes computation possible also creates a possible blind spot.

Continue the Representation & Tokenisation Series

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading