Sound reaches a microphone as continuous pressure variation. An audio model cannot process that continuous world directly: the signal must be sampled, segmented and represented as a finite sequence of workable units.
Those units may be waveform samples, short-time frames, spectrogram patches, continuous learned features, phonetic-like clusters or discrete neural codec tokens. Audio tokenisation is therefore not one fixed algorithm. It is the family of boundary and representation choices that turn time-varying sound into model-addressable structure.
This article extends the eduKateSingapore Representation and Tokenisation series into sound.
The Audio Representation Route
WORLD SOUND → MICROPHONE → ANALOG ELECTRICAL SIGNAL → SAMPLING + QUANTISATION → DIGITAL WAVEFORM → FRAMES / TIME-FREQUENCY REPRESENTATION → LEARNED FEATURES OR DISCRETE AUDIO TOKENS → CONTEXTUAL MODEL → TEXT / AUDIO / ACTION → WORLD RETURN
1. The Microphone Is the First Representation Gate
A microphone converts air-pressure changes into an electrical signal. Frequency response, directionality, distance, room acoustics and background noise already shape what reaches the digital pipeline.
Audio tokenisation therefore begins after physical capture has already selected and distorted part of the acoustic world.
2. Sampling Turns Continuous Time Into Discrete Measurements
Digital audio records the waveform at a finite sampling rate. At 16 kHz, for example, the system stores 16,000 amplitude samples per second.
This is temporal discretisation before any model token exists.
3. Sampling Rate Sets a Frequency Ceiling
Under the sampling theorem, a digital system can represent frequencies only below a limit determined by the sampling rate. Audio captured for speech may use a lower sampling rate than high-fidelity music.
The chosen rate therefore determines which acoustic detail survives upstream.
4. Quantisation Turns Amplitude Into Finite Values
Each sampled amplitude must also be represented with finite numerical precision. Bit depth determines the available quantisation levels.
Sampling divides time; amplitude quantisation divides signal magnitude. Both occur before learned tokenisation.
5. Raw Samples Are Usually Too Fine-Grained for High-Level Modelling
One second of audio can contain tens of thousands of waveform samples. Processing every sample as a high-level token creates extremely long sequences.
Audio models therefore usually construct larger temporal units or compressed features.
6. Frames Create Short Time Windows
Speech and audio analysis often groups consecutive samples into short overlapping frames measured in milliseconds. Within one short frame, aspects of the signal can be treated as approximately stable enough for local analysis.
Frames are the audio analogue of image patches or text chunks: convenient boundaries imposed on a continuous source.
7. Frame Length Controls Temporal Resolution
Short frames preserve rapid temporal change but provide less frequency resolution. Longer frames provide finer frequency information while smoothing over quick events.
The time–frequency trade-off is a representation decision.
8. Frame Overlap Protects Boundary Information
Adjacent audio frames often overlap so a sound event near one frame edge also appears in its neighbour. This reduces sensitivity to arbitrary frame placement.
Overlap improves continuity at the cost of repeated computation—exactly the same boundary trade-off seen in document and image tiling.
9. Window Functions Reduce Artificial Edge Effects
When a finite frame is analysed in frequency space, abrupt cutoffs at frame edges can create spectral artifacts. Window functions taper the frame toward its boundaries.
The representation is deliberately modified to improve the fidelity of the analysis.
10. Spectrograms Convert Time Into Time–Frequency Structure
A spectrogram represents how energy is distributed across frequency over time. It turns one-dimensional waveform samples into a two-dimensional time–frequency image-like representation.
Audio can therefore be tokenised through visual-style patches once converted into a spectrogram.
11. Mel Spectrograms Approximate Human Frequency Resolution
Speech systems often map frequencies onto mel-scaled filter banks so the representation allocates resolution more similarly to aspects of human auditory perception.
This is a receiver-informed transformation: the representation is shaped around useful perceptual distinctions rather than preserving raw frequency uniformly.
12. Spectrogram Pixels Are Not Phonemes
A small time–frequency patch can contain fragments of speech, harmonics or transient noise. Like image patches, these geometric units do not automatically align with meaningful objects.
Higher layers must reconstruct acoustic and linguistic structure across many local units.
13. Speech Has Several Possible Token Scales
- waveform samples;
- short acoustic frames;
- spectral feature vectors;
- phoneme-like units;
- syllable-like units;
- word pieces after speech recognition;
- learned discrete acoustic tokens;
- semantic speech tokens.
Each scale preserves different information.
14. Speech Recognition Creates a Second Tokenisation Layer
An automatic speech recogniser converts acoustic representations into text. The resulting transcript can then be passed through an ordinary text tokenizer.
The system therefore moves from acoustic units to linguistic tokens through an intermediate recognition step.
15. Transcription Removes Some Acoustic Information
A transcript preserves words while often discarding pitch, timing, laughter, hesitation, speaker overlap, emotion and background sound.
Text can be semantically useful while remaining a lossy representation of the audio event.
16. Prosody Is Meaningful Context
Stress, intonation, rhythm and pauses can change interpretation. “Really?” and “Really.” can contain the same lexical word while carrying different pragmatic force.
Audio-native representations can preserve distinctions that transcripts flatten.
17. Speaker Identity Is Another Representation Channel
Who is speaking can matter independently of what is said. Speaker embeddings or diarisation labels can represent identity or speaker turns alongside the acoustic content.
A complete conversational representation may need words, timing and speaker structure together.
18. wav2vec 2.0 Learns From Raw Audio Features
Modern self-supervised speech systems can learn representations directly from waveform-derived features. wav2vec 2.0 learns latent speech representations and uses quantised latent targets during pretraining.
This is an important bridge between continuous acoustic features and discrete learned units.
19. Learned Acoustic Units Need Not Equal Phonemes
A model can discover useful clusters or codebook entries that help prediction without matching a linguist’s phoneme inventory exactly.
As with text tokenizers, computational units can be useful without being canonical linguistic atoms.
20. Neural Audio Codecs Create Discrete Token Streams
Neural codecs encode waveform audio into compressed latent representations and can quantise those latents into discrete codebook indices. Models such as EnCodec demonstrate high-fidelity neural audio compression using learned representations.
These indices can function as audio tokens for generative models.
21. Codec Tokens Preserve More Than Words
Unlike a transcript, acoustic codec tokens can preserve voice quality, timing, background sound and other waveform characteristics depending on bitrate and architecture.
They represent sound rather than only linguistic content.
22. Multiple Codebooks Can Represent Different Detail Levels
Residual vector quantisation can use several codebooks in sequence. Early codebooks capture coarse signal structure; later codebooks refine the reconstruction.
This creates multiple discrete token streams per time step.
23. Bitrate Is a Representation Budget
A higher audio bitrate allows more information to survive compression. Lower bitrate reduces storage and sequence volume but sacrifices fine acoustic detail.
Bitrate therefore plays a role analogous to image resolution or text token budget.
24. Semantic and Acoustic Tokens Solve Different Jobs
A speech model can learn tokens that emphasise linguistic content while an audio codec emphasises reconstructing the waveform. The first may discard speaker timbre; the second may preserve it.
“Audio token” must therefore name which invariants the representation is intended to preserve.
25. Music Needs Different Units From Speech
Music contains pitch, rhythm, harmony, instrumentation and long-term structure. A speech-optimised acoustic representation may not preserve the relationships needed for musical generation or analysis.
Domain-specific tokenisation can operate on notes, events, frames or learned codec units.
26. MIDI Is Already a Symbolic Audio Representation
MIDI does not store sound waveforms. It stores symbolic performance events such as note-on, note-off, pitch, velocity and control changes.
MIDI tokenisation therefore begins from a structured symbolic representation rather than raw audio.
27. Environmental Sound Does Not Have Word Boundaries
A door slam, engine hum, rainstorm and distant conversation can overlap. There is no universal segmentation into one acoustic object at a time.
Audio models need temporal and source-separation representations when multiple sounds coexist.
28. Silence Is Not Necessarily Nothing
Pauses can signal turn-taking, hesitation, emphasis or event boundaries. A pipeline that deletes silence indiscriminately can remove temporal meaning.
Absence of sound can itself be a represented state.
29. Voice Activity Detection Is a Segmentation System
Voice activity detection classifies regions as speech or non-speech. It helps reduce unnecessary processing but can cut off quiet syllables or misclassify background speech.
Its boundaries influence every later acoustic token.
30. Diarisation Creates Speaker Segments
Speaker diarisation answers “who spoke when?” by dividing an audio recording into speaker-labelled regions.
This is higher-level segmentation layered above acoustic frames.
31. Overlapping Speech Breaks Simple Segmentation
Two speakers can talk at the same time. A timeline cannot always be divided into non-overlapping speaker turns without losing the fact that multiple sources coexist.
Representation sometimes needs parallel tracks rather than one linear sequence.
32. Audio Tokens Have Time Duration
A text token has an order position but no natural physical duration. An audio token often corresponds to a time span. Duration, frame hop and alignment therefore become first-class metadata.
Temporal representation cannot be reduced to sequence index alone.
33. Variable-Rate Tokenisation Can Allocate More Units to Complex Regions
Some representations need not use one fixed number of tokens per second. Quiet or predictable regions can be compressed more heavily while complex transients receive more detail.
This is adaptive resolution over time.
34. Audio Generation Needs a Decoder
If a generative model predicts discrete acoustic or codec tokens, a decoder must convert those tokens back into waveform audio.
The final sound quality depends on both token prediction and decoder fidelity.
35. Reconstruction Quality Is Not Semantic Accuracy
A codec can reconstruct a waveform that sounds close to the original while a speech recogniser mishears a word. Conversely, a transcript can be perfectly accurate while losing the speaker’s emotion.
Different audio representations preserve different invariants.
36. Audio Tokens Can Enter Multimodal Models
When audio is processed alongside text, images or video, its token rate determines how much of the shared model context it consumes.
This is explored in Multimodal Tokenisation.
37. Discrete Codebooks Link Audio to General Token Models
Once continuous acoustic features are mapped to finite codebook IDs, the signal becomes a discrete token stream that can be predicted using architectures similar to those used for text.
38. The Audio Tokenisation Audit
- What microphone and sampling rate capture the source?
- What amplitude precision is retained?
- Are waveform samples, frames or spectrograms the working representation?
- What frame length and hop size are used?
- What frequency scale is represented?
- Are silence and background sounds preserved?
- Is speaker identity needed?
- Does the task require words, prosody or full waveform fidelity?
- Are learned units continuous or discrete?
- What bitrate or token rate is used?
- Can overlapping speakers be represented?
- How is time alignment preserved?
- What decoder reconstructs audio?
- Does reconstruction preserve the receiver’s required distinctions?
39. What Students Should Remember
- Sound begins as a continuous physical signal.
- Sampling and quantisation create digital audio before model tokenisation.
- Frames divide time into workable windows.
- Spectrograms represent time and frequency together.
- Transcripts are a lossy linguistic representation of audio.
- Neural codecs can produce discrete acoustic tokens.
- Audio tokenisation must preserve timing.
- Speech, music and environmental sound need different representation priorities.
40. The Deep Principle
Audio tokenisation is the art of turning a continuous event into a finite sequence without erasing the temporal and acoustic distinctions the receiver still needs.
The waveform is not the sound event, the frame is not the word, and the codec token is not the meaning. Each is a boundary chosen so computation can enter a world that never arrived in discrete pieces.
