Music unfolds through pitch, rhythm, duration, dynamics, harmony, articulation and repetition. A model cannot work with “music” as one indivisible object. It needs a representation that turns musical structure into addressable units.
Music tokenisation is the process of deciding whether those units should be notes, MIDI events, time shifts, durations, chords, bars, phrases or learned acoustic tokens. The correct choice depends on whether the receiver needs score generation, expressive performance, analysis, retrieval or raw audio synthesis.
This article continues the eduKateSingapore Representation and Tokenisation series and complements Audio Tokenisation.
The Music Representation Route
MUSICAL IDEA / PERFORMANCE → SCORE OR AUDIO CAPTURE → NOTES / EVENTS / TIMING → MUSIC TOKENS → SEQUENCE MODEL → GENERATED TOKENS → SCORE / MIDI / AUDIO RENDERER → PERFORMANCE OR PLAYBACK → WORLD RETURN
1. Music Has Several Valid Primitive Units
A written score suggests notes and rests. MIDI suggests note-on, note-off, velocity and control events. Audio suggests frames or codec tokens.
Music tokenisation begins by choosing which musical representation is authoritative for the task.
2. A Note Is More Than a Pitch
Pitch alone does not tell us when the note starts, how long it lasts or how loudly it is played.
A complete note representation may need pitch + onset + duration + velocity + instrument.
3. Event-Based Tokenisation Represents Actions in Time
MIDI-like models can use tokens such as NOTE_ON_60, TIME_SHIFT_100ms and NOTE_OFF_60.
The sequence directly records performance events rather than a static score.
4. Time-Shift Tokens Make Silence and Duration Explicit
Instead of attaching absolute timestamps to every event, a model can emit a token indicating how much time passes before the next event.
Relative timing becomes part of the token vocabulary.
5. Relative Timing Matters Deeply in Music
Musical structure depends on interval and temporal relation more than wall-clock position.
Music Transformer highlighted this by using relative attention to model long-range musical structure and timing relationships more effectively.
6. Absolute Time Is Still Useful for Alignment
Scores, stems and multimodal systems may need exact bar, beat or clock positions to align instruments, lyrics and video.
Relative and absolute time solve different jobs.
7. Note-On and Note-Off Preserve Performance Independence
A pianist can press a key, hold it and release it later. Treating duration as a separate event allows overlapping notes naturally.
This matters for chords and polyphony.
8. Duration Tokens Offer a More Compact Alternative
A score-oriented representation can encode pitch plus duration directly rather than separate note-off events.
This shortens sequences but makes overlapping voices more difficult to represent.
9. Polyphony Breaks Simple One-Note-at-a-Time Assumptions
Several notes can sound simultaneously. A strictly linear token stream must choose an arbitrary ordering among events sharing the same onset.
Concurrency is a representation problem, not a musical error.
10. Chords Are Higher-Level Music Tokens
C major can be represented as three simultaneous pitches or as one chord label.
The chord token compresses several notes into one harmonic unit.
11. Chord Labels Lose Voicing Detail
C major does not specify inversion, spacing or octave placement.
Coarse harmonic tokens preserve function while discarding performance detail.
12. Rhythm Is More Than Timestamp Differences
Musical rhythm relates events to beats, metre and recurring patterns.
Beat-relative tokens can be more musically meaningful than raw milliseconds.
13. Quantisation Snaps Performance to a Grid
A human performance may place notes slightly ahead of or behind the beat. Quantising to sixteenth notes makes the score cleaner while removing expressive timing.
Quantisation is a deliberate loss of microtiming.
14. Fine Rhythmic Grids Preserve More Expression
Smaller time divisions retain subtle timing but enlarge the timing vocabulary or sequence length.
Temporal resolution has the same cost trade-off as spatial and textual granularity.
15. Tempo Converts Beat Time Into Clock Time
The same rhythmic pattern can be performed at different tempos.
Beat-relative structure and absolute duration should remain separable.
16. Tempo Changes Can Be Tokens
MIDI files can include tempo events. A model can therefore represent acceleration and slowing explicitly rather than infer them from timestamps alone.
17. Velocity Tokens Represent Dynamics
MIDI velocity approximates attack intensity. Quantising velocity into bins gives a finite dynamics vocabulary.
Too few bins flatten expressive contrast; too many create sparse labels.
18. Dynamics Symbols Are Higher-Level Score Tokens
p, f, crescendo and diminuendo describe expressive intent rather than one exact velocity value.
Notation and performance encode different layers of the same musical idea.
19. Instruments Need Identity
The same pitch played by violin and piano has different timbre.
Instrument tokens preserve source identity in symbolic arrangements.
20. Multi-Track Music Adds Parallel Sequences
An orchestration contains several simultaneous instrument streams.
Models can interleave events into one sequence or preserve track identity explicitly.
21. Track Order Should Not Become Accidental Meaning
Reordering two independent MIDI tracks should not change the composition’s core musical content.
Track identity matters more than arbitrary serialization order.
22. Bars and Measures Are Higher-Level Chunks
A bar groups events under a metre and beat cycle.
Bar tokens or markers help models learn periodic structure and phrase boundaries.
23. Phrases Are Coarser Musical Tokens
A musical phrase can span several bars and function like a sentence.
Phrase-level representation helps long-range planning but requires reliable segmentation.
24. Motifs Are Reusable Pattern Tokens
A motif can recur in transposed, inverted or rhythmically altered form.
Exact token repetition is therefore too narrow a definition of musical recurrence.
25. Transposition Changes Pitch Tokens While Preserving Relations
A melody moved up a whole tone has different absolute note tokens and similar interval structure.
Relative pitch representations can preserve transposition invariance more directly.
26. Interval Tokens Represent Relationships Between Notes
Instead of encoding C then E, a model can encode a starting pitch plus an upward major-third interval.
Relative representation emphasizes melodic contour over absolute register.
27. Key Is Context, Not One Note
The same pitch can serve different harmonic functions in different keys.
Key tokens provide global tonal context above local note identities.
28. Tonal Music and Atonal Music Need Different Assumptions
A representation centred on key and chord functions can be powerful for tonal music and poorly matched to music organised by other principles.
Representation should not force one musical theory onto every repertoire.
29. MIDI Is Symbolic, Not Audio
MIDI specifies musical events and controls. It does not store microphone waveforms.
The same MIDI performance can sound radically different under another instrument or synthesiser.
30. Audio Tokens Preserve Timbre Better
Neural audio codecs can represent the actual acoustic signal, including timbre, room and performance nuance.
Symbolic music tokens and audio tokens preserve different invariants.
31. Score Tokens Preserve Compositional Structure Better
A score makes pitch, rhythm, voices and notation explicit.
It is easier to edit harmonically and structurally than a raw waveform.
32. Lyrics Add Language Tokens
Songs can align text syllables with note events.
This creates a multimodal alignment problem between linguistic and musical sequences.
33. Syllable-to-Note Alignment Is Not Always One-to-One
One syllable can span several notes; several syllables can fit one rapid rhythmic group.
Alignment needs duration and grouping information.
34. Long-Range Repetition Is Central to Musical Coherence
Verses, choruses, themes and returns can recur minutes apart.
Context length determines whether a model can directly compare distant musical sections.
35. More Context Does Not Guarantee Better Composition
A model can remember earlier tokens and still fail to understand formal function or development.
Sequence memory and musical planning are distinct capabilities.
36. Generated MIDI Needs Structural Validation
A syntactically valid MIDI file can contain impossible instrument ranges, excessive density or awkward timing.
Format validity is not musical quality.
37. Human Performance Is the Final Receiver Test
A generated sequence can look coherent in tokens and feel unplayable or lifeless to musicians.
World return means listening, performing and evaluating the music as music.
38. The Music Tokenisation Audit
- Is the source a score, MIDI performance or audio waveform?
- What counts as one token: note, event, duration, chord, bar or phrase?
- How is time represented?
- How fine is rhythmic quantisation?
- Are velocity and articulation preserved?
- How are simultaneous notes serialized?
- How is instrument identity represented?
- Are key, metre and tempo explicit?
- Can motifs and transpositions be represented relationally?
- What long-range context is available?
- Are score and performance representations distinguished?
- Can lyrics align with musical time?
- Can generated tokens reconstruct a valid score or MIDI file?
- Does the result survive listening and performance?
39. What Students Should Remember
- Music can be tokenised as notes, events, durations, chords or larger phrases.
- Relative time is central to musical structure.
- Quantisation simplifies rhythm while losing microtiming.
- MIDI is symbolic representation, not audio.
- Chords and phrases are higher-level tokens.
- Relative pitch can preserve transposition structure.
- Long-range repetition matters for musical coherence.
- The final receiver test is musical, not merely syntactic.
40. The Deep Principle
Music tokenisation makes temporal relationships addressable. It works when the sequence preserves not merely which notes occurred, but how timing, hierarchy and repetition turn those notes into music.
A note token names an event. Musical meaning appears when the model preserves the relations among events across beat, phrase, harmony and time.