Tokenisation is the boundary-making stage that turns an input representation into a sequence of units a computational system can address. It sounds like a small preprocessing step. In practice, it determines vocabulary size, sequence length, coverage, reconstruction behaviour, multilingual efficiency and part of the computational geometry that every later layer inherits.
This article sits underneath Representation and Tokenisation | How Information Becomes Units a System Can Work With and under the canonical World Representation & Cognitive Tools owner. The job here is narrower: follow the mechanism from source string to workable units and back again.
The Pipeline in One Line
SOURCE TEXT → ENCODING → NORMALISATION → PRE-TOKENISATION → SEGMENTATION → TOKEN VOCABULARY LOOKUP → TOKEN IDs → MODEL INPUT → OUTPUT IDs → DECODING → RECONSTRUCTED TEXT
Different tokenisers rearrange, combine or omit some stages, but this sequence is a useful control model. It lets us ask where a behaviour came from instead of blaming everything on the language model itself.
1. Begin With the Representation You Actually Have
A tokenizer does not receive meaning directly. It receives a representation: usually an encoded text string, sometimes after another layer has already extracted text from speech, images, documents or structured data. If optical character recognition mistakes a letter, the tokenizer faithfully tokenises the mistake. If a PDF extractor loses column order, the tokenizer receives the wrong sequence. If a transcription system drops a speaker label, tokenisation cannot restore that lost identity by itself.
Tokenisation must therefore be analysed as one stage in a longer chain. A clean tokenizer cannot repair an unclean source representation unless some separate correction mechanism exists.
2. Encoding: Characters Must Become Machine-Representable
Modern text systems typically work with Unicode, which assigns code points to characters and symbols across writing systems. But the same visible character sequence can sometimes have more than one underlying code-point sequence. This is why text that looks identical to a reader may fail a byte-level or code-point-level equality test.
Unicode normalisation addresses classes of equivalent representations. The Unicode Consortium specifies standard normalisation forms in Unicode Standard Annex #15. The lesson for tokenisation is not that every system should normalise everything identically. It is that the representation entering the tokenizer is already a technical choice.
3. Normalisation: Decide Which Differences Should Survive
Normalisation can lowercase text, standardise equivalent Unicode sequences, canonicalise whitespace or perform other transformations. Each operation changes what distinctions remain available. Lowercasing can improve vocabulary sharing across sentence positions, but it also removes a distinction that may matter for proper names, acronyms or case-sensitive code. Whitespace cleanup can simplify text, but it may destroy indentation that carries meaning in programming languages or formatted documents.
The correct question is not whether normalisation is good. It is: which distinctions are safe to collapse for this receiver and task?
4. Pre-Tokenisation: Establish Candidate Regions
Some tokenizer pipelines first split or mark the input around spaces, punctuation or other patterns before applying a learned subword model. Others treat raw text more directly. Pre-tokenisation can influence every downstream boundary because it may prohibit a learned token from crossing particular separators.
This is one reason tokenisers that appear to use the same broad algorithm can still produce different segmentations. The complete pipeline matters, not only the label BPE, WordPiece or Unigram.
5. The Vocabulary Is a Finite Set of Available Pieces
A tokenizer vocabulary is a controlled inventory of pieces the system knows how to identify. It can contain whole words, common fragments, punctuation-associated forms, bytes, characters and special control symbols. A larger vocabulary can represent frequent patterns compactly but increases the number of available token identities. A smaller vocabulary is easier to cover but tends to produce longer sequences.
LARGER VOCABULARY → MORE COMPACT COMMON PATTERNS → MORE TOKEN TYPES SMALLER VOCABULARY → BROADER COMPOSITION FROM SMALL PIECES → LONGER SEQUENCES
6. Why Whole-Word Tokenisation Breaks at Scale
Whole-word tokenisation is attractive because human readers already recognise words. But language is productive. New names, compounds, inflections, spelling variants and specialist terms appear continuously. A word vocabulary large enough to include everything becomes impractical and still cannot guarantee coverage of tomorrow’s vocabulary.
The alternative is compositionality: represent uncommon words through smaller reusable parts. Subword tokenisation became a standard solution because it combines useful compression with broad coverage.
7. Byte Pair Encoding: Build Frequent Larger Pieces
Byte Pair Encoding, adapted for subword tokenisation, starts from small units and repeatedly combines frequent neighbouring pairs according to a learned merge process. Common strings can become single vocabulary items while rare strings remain decomposable into smaller units.
Rico Sennrich, Barry Haddow and Alexandra Birch’s 2016 paper Neural Machine Translation of Rare Words with Subword Units helped establish the value of this approach for open-vocabulary neural language processing. The key systems insight is that frequency can be used to choose which boundaries deserve compression.
BPE-style segmentation is useful, but it does not claim that every learned piece is a linguistic morpheme. A frequent character sequence may be computationally useful even when it has no clean grammatical status.
8. WordPiece: Another Route to Reusable Subwords
WordPiece is another influential subword family associated with BERT-style systems. It also builds a subword vocabulary, but its training and segmentation logic differs from standard BPE. Practical explanations are available in the Hugging Face tokenisation documentation.
The important educational point is that two subword tokenisers can receive the same visible text and choose different boundaries because their vocabularies and segmentation rules differ.
9. Unigram: Optimise a Probabilistic Vocabulary
Unigram tokenisation takes a probabilistic approach. It can begin with a large candidate inventory and estimate which pieces contribute most to representing the training data. Segmentation is selected under a learned probability model. This makes the underlying principle visible: tokenisation is optimisation under a vocabulary and objective, not discovery of one inevitable cut.
10. SentencePiece: Raw Text as the Starting Point
Taku Kudo and John Richardson described SentencePiece as a language-independent subword tokenizer and detokenizer that can train directly from raw sentences. This matters because many older pipelines assumed that whitespace-delimited words already existed as a preprocessing layer, an assumption that does not transfer cleanly across all writing systems.
11. Byte-Level Routes: Coverage From a Lower Layer
Byte-level approaches can provide broad coverage because encoded text can ultimately be expressed through bytes. This reduces dependence on a fixed character inventory and gives a route for unusual strings, mixed scripts and arbitrary text. The trade-off is that the pieces may be less intuitive from a human linguistic perspective.
12. Special Tokens Are Part of the Protocol
Many model pipelines reserve special tokens or identifiers for sequence boundaries, padding, masks, roles, separators or control states. These units may not correspond to visible source text at all. They exist because the tokenizer-model pair is also a communication protocol.
13. Tokens Become IDs
Once pieces are chosen, each token maps to an identifier. The identifier is usually an integer indexing the vocabulary. If token 4172 corresponds to a particular piece, the number 4172 does not contain the meaning of that piece. It is an address.
This matters because adjacent token IDs do not normally imply semantic similarity. Semantic structure is learned in later numerical representations, not in the arbitrary ordering of vocabulary IDs.
14. IDs Become Numerical Representations
The model maps token identities into numerical vectors or other internal representations suitable for learned computation. Position or ordering information is also incorporated because the sequence dog bites man is not equivalent to man bites dog.
15. Context Changes the Working Meaning
A token’s initial identity does not settle its final interpretation. Contextual sequence models repeatedly transform internal states using information from surrounding positions. The token representing cell behaves differently in discussions of biology, prisons, spreadsheets and batteries because the sequence around it alters the working representation.
16. Decoding Is the Return Path
A tokenizer usually supports decoding: converting token identities back into human-visible text. A well-designed tokenizer can make this route deterministic for ordinary inputs, but special tokens, normalisation choices or invalid sequences complicate the idea of perfect reconstruction.
17. Reversibility Is Not Semantic Fidelity
A tokenizer can reconstruct every character and still choose boundaries that are inconvenient for a downstream task. Reversibility answers whether the source can return. Semantic utility answers whether the chosen pieces support useful computation. These are different tests.
18. Whitespace Is Not Empty
Human readers often treat spaces as absence. Tokenizers may treat them as meaningful boundary signals or incorporate them into token pieces. In code, indentation can be structural. In poetry, spacing can carry form. In tables and fixed-width formats, whitespace can represent columns. Casual cleanup can therefore destroy information before tokenisation begins.
19. Punctuation Is Not Decorative
Punctuation can mark sentence structure, quotation, mathematical grouping, paths, URLs, programming syntax or emotive tone. A tokenizer may isolate punctuation, attach it to neighbouring text or learn multi-character combinations. In ordinary prose a comma may be easy to infer; in a regular expression or programming language, one mark can change the entire object.
20. Numbers Reveal the Difference Between Symbol and Concept
Humans learn that 2048 is one number with place-value structure. A tokenizer may represent the digit string using one piece or several. The model must therefore learn mathematical relationships through a representation that may not align with the conceptual decomposition taught in mathematics.
21. Code Has Different Natural Boundaries
Programming languages contain identifiers, operators, indentation, delimiters, literals and punctuation-dense expressions. A tokenizer built mainly from ordinary prose may fragment code differently from one trained with substantial code. The same principle applies to mathematics, chemistry, law and medicine: domain frequency changes which sequences become economical units.
22. Proper Names and Rare Terms Expose Coverage
A rare surname, place name or scientific term may be divided into several familiar pieces. This is not automatically failure. Subword systems are designed to represent unseen forms compositionally. But longer fragmentations can make exact copying, spelling-sensitive generation or entity matching more demanding.
23. Multilingual Tokenisation Is an Allocation Problem
A finite vocabulary allocates capacity across writing systems and patterns seen during tokenizer training. Languages with abundant representation may gain many compact frequent pieces. Others may require longer decompositions for similar amounts of human-visible text. Equal token budgets therefore do not always mean equal amounts of linguistic content.
24. Context Windows Are Measured After Tokenisation
Model context capacity is defined in model-side units, not human words or pages. A densely tokenised language, unusual codebase or punctuation-heavy dataset can consume context faster than ordinary prose. Capacity planning should therefore count with the actual tokenizer rather than rely on a fixed words-per-token rule.
25. Retrieval Adds Another Boundary Layer
Long-document retrieval often divides sources into chunks before embedding, indexing or model ingestion. Chunking is a second boundary system layered above tokenizer boundaries. A bad retrieval cut can separate a definition from its exception, a table heading from its rows or a condition from its qualifier. Tokenisation belongs to a wider family of segmentation problems.
26. Token Counts Belong to a Tokenizer-Text Pair
There is no intrinsic token count attached to a sentence independent of a tokenizer. Change the tokenizer and the count can change. Change Unicode representation, whitespace or punctuation and the segmentation may change again.
A token count is not a property of language alone. It is a property of an input representation processed by a particular tokenizer.
27. Training-Time and Inference-Time Compatibility Matters
A model learns on sequences produced by a particular tokenisation scheme. Using an incompatible vocabulary at inference can break the mapping between token IDs and learned parameters. The tokenizer and model are therefore not freely interchangeable accessories. They are part of one trained interface contract.
28. Version the Tokenizer
When reproducibility matters, record the tokenizer family, vocabulary and exact version or artifact. If a later system silently changes normalisation, special-token handling or vocabulary, historical token counts and model inputs may no longer be identical. Tokenizers belong in production version control just as other interfaces do.
29. Preserve Offsets Back to the Source
Many tokenizer APIs return offsets mapping tokens to character spans in source text. These are valuable for highlighting, annotation, entity tasks and debugging. Offsets make a general systems principle explicit: transformations are safer when they preserve a traceable route back to the source.
30. Test the Round Trip
A basic QA test is encode → decode → compare. The correct comparison depends on normalisation policy. If equivalent forms are intentionally collapsed, exact byte equality may not be the right invariant. The test must first declare what is required to survive.
31. Test Adversarial Boundaries
Do not test only ordinary prose. Test long numbers, mixed scripts, emoji sequences, uncommon names, repeated punctuation, zero-width characters, unusual whitespace, code, URLs, equations and malformed input. These cases reveal assumptions hidden by familiar sentences.
32. Do Not Infer Linguistic Truth From Token Boundaries
If a tokenizer splits a word into three pieces, that does not prove the language contains three morphemes there. If it keeps a phrase fragment together, that does not prove the fragment is one semantic unit. The tokenizer optimises a computational objective under a particular corpus and vocabulary budget.
33. Do Not Infer Understanding From Compact Tokens
A common word may be represented by one token while a rare technical term uses several. That says something about representation efficiency, not necessarily conceptual competence. Tokenisation is one constraint in the system, not a complete explanation of behaviour.
34. Human Expertise Also Changes Effective Tokens
Experts chunk information. A chess expert sees a formation where a novice sees isolated pieces. A mathematician sees a known structure where a beginner sees symbols. A fluent reader sees phrases where an early reader decodes letters. Good units reduce working-memory load while preserving structure needed for reasoning.
35. Tokenisation Debugging Checklist
- What exact source representation enters the tokenizer?
- What encoding and normalisation rules apply?
- Is there pre-tokenisation?
- Which vocabulary and tokenizer version are used?
- Which algorithm family performs segmentation?
- How are spaces, punctuation and special symbols handled?
- How are rare sequences represented?
- Can tokens map back to source offsets?
- Can encoding and decoding preserve the required invariants?
- How many tokens do representative languages and domains consume?
- What happens to long numbers, code, URLs and unusual scripts?
- Is the tokenizer exactly compatible with the model?
36. The Deeper Principle
Tokenisation is not merely text chopping. It is a negotiated boundary between human-visible representation and model-ready computation. The tokenizer decides which recurring forms deserve compact identities, which rare forms must be composed from smaller pieces and which differences are preserved or collapsed before the model begins its main work.
Before intelligence can operate on a world, the world must be represented. Before a representation can be computed over, it must be given workable structure. Every boundary that makes computation possible also creates a possible blind spot.
