Audio Tokenisation | How Sound Becomes Frames, Features and Discrete Acoustic Units

Sound reaches a microphone as continuous pressure variation. An audio model cannot process that continuous world directly: the signal must be sampled, segmented and represented as a finite sequence of workable units.

Those units may be waveform samples, short-time frames, spectrogram patches, continuous learned features, phonetic-like clusters or discrete neural codec tokens. Audio tokenisation is therefore not one fixed algorithm. It is the family of boundary and representation choices that turn time-varying sound into model-addressable structure.

This article extends the eduKateSingapore Representation and Tokenisation series into sound.

The Audio Representation Route

WORLD SOUND
→ MICROPHONE
→ ANALOG ELECTRICAL SIGNAL
→ SAMPLING + QUANTISATION
→ DIGITAL WAVEFORM
→ FRAMES / TIME-FREQUENCY REPRESENTATION
→ LEARNED FEATURES OR DISCRETE AUDIO TOKENS
→ CONTEXTUAL MODEL
→ TEXT / AUDIO / ACTION
→ WORLD RETURN

1. The Microphone Is the First Representation Gate

A microphone converts air-pressure changes into an electrical signal. Frequency response, directionality, distance, room acoustics and background noise already shape what reaches the digital pipeline.

Audio tokenisation therefore begins after physical capture has already selected and distorted part of the acoustic world.

2. Sampling Turns Continuous Time Into Discrete Measurements

Digital audio records the waveform at a finite sampling rate. At 16 kHz, for example, the system stores 16,000 amplitude samples per second.

This is temporal discretisation before any model token exists.

3. Sampling Rate Sets a Frequency Ceiling

Under the sampling theorem, a digital system can represent frequencies only below a limit determined by the sampling rate. Audio captured for speech may use a lower sampling rate than high-fidelity music.

The chosen rate therefore determines which acoustic detail survives upstream.

4. Quantisation Turns Amplitude Into Finite Values

Each sampled amplitude must also be represented with finite numerical precision. Bit depth determines the available quantisation levels.

Sampling divides time; amplitude quantisation divides signal magnitude. Both occur before learned tokenisation.

5. Raw Samples Are Usually Too Fine-Grained for High-Level Modelling

One second of audio can contain tens of thousands of waveform samples. Processing every sample as a high-level token creates extremely long sequences.

Audio models therefore usually construct larger temporal units or compressed features.

6. Frames Create Short Time Windows

Speech and audio analysis often groups consecutive samples into short overlapping frames measured in milliseconds. Within one short frame, aspects of the signal can be treated as approximately stable enough for local analysis.

Frames are the audio analogue of image patches or text chunks: convenient boundaries imposed on a continuous source.

7. Frame Length Controls Temporal Resolution

Short frames preserve rapid temporal change but provide less frequency resolution. Longer frames provide finer frequency information while smoothing over quick events.

The time–frequency trade-off is a representation decision.

8. Frame Overlap Protects Boundary Information

Adjacent audio frames often overlap so a sound event near one frame edge also appears in its neighbour. This reduces sensitivity to arbitrary frame placement.

Overlap improves continuity at the cost of repeated computation—exactly the same boundary trade-off seen in document and image tiling.

9. Window Functions Reduce Artificial Edge Effects

When a finite frame is analysed in frequency space, abrupt cutoffs at frame edges can create spectral artifacts. Window functions taper the frame toward its boundaries.

The representation is deliberately modified to improve the fidelity of the analysis.

10. Spectrograms Convert Time Into Time–Frequency Structure

A spectrogram represents how energy is distributed across frequency over time. It turns one-dimensional waveform samples into a two-dimensional time–frequency image-like representation.

Audio can therefore be tokenised through visual-style patches once converted into a spectrogram.

11. Mel Spectrograms Approximate Human Frequency Resolution

Speech systems often map frequencies onto mel-scaled filter banks so the representation allocates resolution more similarly to aspects of human auditory perception.

This is a receiver-informed transformation: the representation is shaped around useful perceptual distinctions rather than preserving raw frequency uniformly.

12. Spectrogram Pixels Are Not Phonemes

A small time–frequency patch can contain fragments of speech, harmonics or transient noise. Like image patches, these geometric units do not automatically align with meaningful objects.

Higher layers must reconstruct acoustic and linguistic structure across many local units.

13. Speech Has Several Possible Token Scales

Each scale preserves different information.

14. Speech Recognition Creates a Second Tokenisation Layer

An automatic speech recogniser converts acoustic representations into text. The resulting transcript can then be passed through an ordinary text tokenizer.

The system therefore moves from acoustic units to linguistic tokens through an intermediate recognition step.

15. Transcription Removes Some Acoustic Information

A transcript preserves words while often discarding pitch, timing, laughter, hesitation, speaker overlap, emotion and background sound.

Text can be semantically useful while remaining a lossy representation of the audio event.

16. Prosody Is Meaningful Context

Stress, intonation, rhythm and pauses can change interpretation. “Really?” and “Really.” can contain the same lexical word while carrying different pragmatic force.

Audio-native representations can preserve distinctions that transcripts flatten.

17. Speaker Identity Is Another Representation Channel

Who is speaking can matter independently of what is said. Speaker embeddings or diarisation labels can represent identity or speaker turns alongside the acoustic content.

A complete conversational representation may need words, timing and speaker structure together.

18. wav2vec 2.0 Learns From Raw Audio Features

Modern self-supervised speech systems can learn representations directly from waveform-derived features. wav2vec 2.0 learns latent speech representations and uses quantised latent targets during pretraining.

This is an important bridge between continuous acoustic features and discrete learned units.

19. Learned Acoustic Units Need Not Equal Phonemes

A model can discover useful clusters or codebook entries that help prediction without matching a linguist’s phoneme inventory exactly.

As with text tokenizers, computational units can be useful without being canonical linguistic atoms.

20. Neural Audio Codecs Create Discrete Token Streams

Neural codecs encode waveform audio into compressed latent representations and can quantise those latents into discrete codebook indices. Models such as EnCodec demonstrate high-fidelity neural audio compression using learned representations.

These indices can function as audio tokens for generative models.

21. Codec Tokens Preserve More Than Words

Unlike a transcript, acoustic codec tokens can preserve voice quality, timing, background sound and other waveform characteristics depending on bitrate and architecture.

They represent sound rather than only linguistic content.

22. Multiple Codebooks Can Represent Different Detail Levels

Residual vector quantisation can use several codebooks in sequence. Early codebooks capture coarse signal structure; later codebooks refine the reconstruction.

This creates multiple discrete token streams per time step.

23. Bitrate Is a Representation Budget

A higher audio bitrate allows more information to survive compression. Lower bitrate reduces storage and sequence volume but sacrifices fine acoustic detail.

Bitrate therefore plays a role analogous to image resolution or text token budget.

24. Semantic and Acoustic Tokens Solve Different Jobs

A speech model can learn tokens that emphasise linguistic content while an audio codec emphasises reconstructing the waveform. The first may discard speaker timbre; the second may preserve it.

“Audio token” must therefore name which invariants the representation is intended to preserve.

25. Music Needs Different Units From Speech

Music contains pitch, rhythm, harmony, instrumentation and long-term structure. A speech-optimised acoustic representation may not preserve the relationships needed for musical generation or analysis.

Domain-specific tokenisation can operate on notes, events, frames or learned codec units.

26. MIDI Is Already a Symbolic Audio Representation

MIDI does not store sound waveforms. It stores symbolic performance events such as note-on, note-off, pitch, velocity and control changes.

MIDI tokenisation therefore begins from a structured symbolic representation rather than raw audio.

27. Environmental Sound Does Not Have Word Boundaries

A door slam, engine hum, rainstorm and distant conversation can overlap. There is no universal segmentation into one acoustic object at a time.

Audio models need temporal and source-separation representations when multiple sounds coexist.

28. Silence Is Not Necessarily Nothing

Pauses can signal turn-taking, hesitation, emphasis or event boundaries. A pipeline that deletes silence indiscriminately can remove temporal meaning.

Absence of sound can itself be a represented state.

29. Voice Activity Detection Is a Segmentation System

Voice activity detection classifies regions as speech or non-speech. It helps reduce unnecessary processing but can cut off quiet syllables or misclassify background speech.

Its boundaries influence every later acoustic token.

30. Diarisation Creates Speaker Segments

Speaker diarisation answers “who spoke when?” by dividing an audio recording into speaker-labelled regions.

This is higher-level segmentation layered above acoustic frames.

31. Overlapping Speech Breaks Simple Segmentation

Two speakers can talk at the same time. A timeline cannot always be divided into non-overlapping speaker turns without losing the fact that multiple sources coexist.

Representation sometimes needs parallel tracks rather than one linear sequence.

32. Audio Tokens Have Time Duration

A text token has an order position but no natural physical duration. An audio token often corresponds to a time span. Duration, frame hop and alignment therefore become first-class metadata.

Temporal representation cannot be reduced to sequence index alone.

33. Variable-Rate Tokenisation Can Allocate More Units to Complex Regions

Some representations need not use one fixed number of tokens per second. Quiet or predictable regions can be compressed more heavily while complex transients receive more detail.

This is adaptive resolution over time.

34. Audio Generation Needs a Decoder

If a generative model predicts discrete acoustic or codec tokens, a decoder must convert those tokens back into waveform audio.

The final sound quality depends on both token prediction and decoder fidelity.

35. Reconstruction Quality Is Not Semantic Accuracy

A codec can reconstruct a waveform that sounds close to the original while a speech recogniser mishears a word. Conversely, a transcript can be perfectly accurate while losing the speaker’s emotion.

Different audio representations preserve different invariants.

36. Audio Tokens Can Enter Multimodal Models

When audio is processed alongside text, images or video, its token rate determines how much of the shared model context it consumes.

This is explored in Multimodal Tokenisation.

37. Discrete Codebooks Link Audio to General Token Models

Once continuous acoustic features are mapped to finite codebook IDs, the signal becomes a discrete token stream that can be predicted using architectures similar to those used for text.

See Discrete Latent Tokens.

38. The Audio Tokenisation Audit

  1. What microphone and sampling rate capture the source?
  2. What amplitude precision is retained?
  3. Are waveform samples, frames or spectrograms the working representation?
  4. What frame length and hop size are used?
  5. What frequency scale is represented?
  6. Are silence and background sounds preserved?
  7. Is speaker identity needed?
  8. Does the task require words, prosody or full waveform fidelity?
  9. Are learned units continuous or discrete?
  10. What bitrate or token rate is used?
  11. Can overlapping speakers be represented?
  12. How is time alignment preserved?
  13. What decoder reconstructs audio?
  14. Does reconstruction preserve the receiver’s required distinctions?

39. What Students Should Remember

40. The Deep Principle

Audio tokenisation is the art of turning a continuous event into a finite sequence without erasing the temporal and acoustic distinctions the receiver still needs.

The waveform is not the sound event, the frame is not the word, and the codec token is not the meaning. Each is a boundary chosen so computation can enter a world that never arrived in discrete pieces.

Continue the Representation & Tokenisation Series

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading