Attention remembers by keeping many earlier token representations available for comparison. State-space models take a different route: they carry sequence history forward through a recurrent internal state.
State-space representation is the problem of deciding what that state should preserve, what it should forget, how new input should update it and whether the resulting compressed history still contains the distinctions needed later.
This article continues the eduKateSingapore Representation and Tokenisation series beyond token creation, position and attention topology into recurrent sequence state.
The State-Space Route
TOKEN SEQUENCE → INPUT-DEPENDENT UPDATE → RECURRENT LATENT STATE → RETAIN / FORGET / TRANSFORM HISTORY → OUTPUT REPRESENTATION → NEXT TOKEN / CLASSIFICATION / CONTROL STATE AT TIME t = COMPRESSED HISTORY OF WHAT THE MODEL CHOSE TO CARRY FORWARD
1. Recurrent State Is a Compressed History
A recurrent sequence model does not necessarily keep every prior token representation explicitly available. Instead, it updates a state vector as new inputs arrive.
The state becomes a rolling summary of the past.
2. State Is Not the Same as Memory Archive
An archive preserves recoverable past records. A recurrent state preserves whatever transformed information the update rule carries forward.
Once information disappears from state, exact recovery may be impossible.
3. State-Space Models Come From Dynamical Systems
Classical state-space systems represent how a hidden state evolves over time under inputs and how outputs are generated from that state.
Modern neural SSMs adapt this idea to sequence modelling.
4. The Hidden State Is a Latent Representation
It need not correspond directly to words, events or human concepts.
Its job is operational: retain enough information about the sequence history to support future prediction or task decisions.
5. Linear State-Space Updates Can Be Highly Efficient
Unlike dense self-attention, which compares many token pairs, recurrent state updates can process a sequence with cost that scales linearly with sequence length.
This makes state-space methods attractive for long sequences.
6. Efficiency Alone Is Not Enough
A model can process a million tokens cheaply and still fail if the recurrent state forgets the one detail needed near the end.
Representation fidelity is more important than nominal sequence capacity.
7. Early SSMs Struggled With Content-Dependent Selection
Traditional linear state-space dynamics apply similar transition rules regardless of the specific token content.
Language and other discrete sequences often require selective memory: some tokens should strongly update state while others should be largely ignored.
8. Mamba Makes State Updates Input-Selective
Mamba: Linear-Time Sequence Modeling with Selective State Spaces makes key SSM parameters depend on the current input, allowing the model to decide what information to propagate or forget based on token content.
This turns state update into a learned content-sensitive filtering process.
9. Selectivity Is a Memory Gate
Not every token deserves equal influence on future state.
A selective state mechanism can amplify, suppress or transform incoming information before it enters long-term recurrent representation.
10. Selectivity Is Related to Adaptive Tokenisation
Adaptive tokenisation decides which representational units deserve continued compute. Selective state-space models decide which incoming information deserves persistent influence.
Both allocate representational life selectively.
11. Forgetting Is Built Into Recurrent State
Finite state cannot preserve every past distinction perfectly.
The update dynamics therefore implement an implicit forgetting policy even when no explicit memory-deletion operation exists.
12. Useful State Must Forget Redundancy Before Exceptions
Repeated formatting or predictable filler can be compressed aggressively.
A rare exception, negation or identity change may need persistent representation far into the future.
13. Long-Range Dependency Is a State-Retention Test
If a name introduced at token 500 determines the answer at token 50,000, the model must preserve the relevant identity through thousands of updates.
Long context is therefore a test of durable state, not just input acceptance.
14. Recurrent State Does Not Provide Random Access Automatically
Attention can directly revisit an earlier token representation if it remains in context.
A compressed recurrent state generally cannot “jump back” to exact earlier content unless that information was preserved in state or stored elsewhere.
15. This Changes the Meaning of Retrieval Inside the Sequence
An attention model can query explicit past token states. A state-space model queries its current summary of the past.
The distinction is direct memory versus compressed memory.
16. Compression Can Improve Robustness by Removing Distractors
A recurrent state that forgets irrelevant surface detail can focus on persistent structure.
The same compression can become harmful when the “surface detail” later turns out to matter.
17. State Dimension Is a Capacity Budget
A larger latent state can carry more information and costs more memory and compute.
A smaller state forces stronger compression and greater competition among remembered features.
18. State Capacity Should Match Task Complexity
Simple local pattern recognition may need little state. Long-range program analysis, genomics or document reasoning may require substantially richer state.
There is no universal state size independent of receiver job.
19. State Update Frequency Matters
Updating at every primitive byte, token or frame gives fine temporal resolution.
Updating over chunks or pooled representations reduces sequence length while shifting detail into the chunk encoder.
20. Hierarchical State Can Operate at Several Timescales
Fast state can capture local syntax or motion; slower state can track topic, scene or process stage.
Multi-timescale recurrence mirrors Hierarchical Representation.
21. State-Space Models Can Process Multiple Modalities
Mamba reports strong sequence-model performance across language, audio and genomics.
The common abstraction is a sequence whose useful history can be carried through selective state.
22. Modality Still Determines What State Must Preserve
Audio needs temporal phase and phonetic structure; genomics needs motif and long-range regulatory context; language needs entities, syntax and discourse.
One architecture can span modalities without making their representational jobs identical.
23. State-Space Models Change Interaction Topology
Information does not move through arbitrary token-to-token attention edges.
It moves forward through state transitions, creating a sequential communication topology.
24. This Is a Different Path Length
Two distant tokens can influence each other only through the chain of intervening state updates.
Selective dynamics must preserve the relevant signal over that path.
25. Bidirectional Use Requires Two Directions or Another Design
A left-to-right recurrent state naturally uses past context. Whole-sequence representation tasks can process both directions or combine recurrent passes.
Directionality should match the information available at inference time.
26. Streaming Is a Natural Strength
Because state can be updated incrementally, a state-space model can process streaming audio, sensor data or long text without storing an ever-growing explicit history.
The representation is naturally online.
27. Streaming Makes Correction Harder
If an earlier input is later discovered to be wrong, a recurrent state may need replay from that point or an explicit correction mechanism.
Compressed history is efficient and less editable.
28. External Retrieval Complements Recurrent State
Facts that are too important to risk forgetting can live in an external store and be retrieved when needed.
State carries working history; retrieval restores addressable source detail.
29. Hybrid Architectures Can Combine Attention and State Space
Attention is strong at direct content-based interaction; recurrent state is strong at efficient sequential compression.
Architectures can mix these mechanisms rather than treat them as mutually exclusive.
30. State Interpretability Is Difficult
A recurrent vector does not come with labels saying which dimensions represent topic, identity or timing.
Probing can reveal information encoded in state without proving the model uses that information causally.
31. State Failure Is Often Invisible Until a Later Query
The model may appear coherent for thousands of steps after silently forgetting a critical premise.
Evaluation should place decisive information far in the past and test whether it remains usable.
32. Sequence Length Benchmarks Should Test Content Retention
Processing a million positions is not equivalent to remembering useful information across a million positions.
Needle retrieval, multi-step dependencies and long-range state changes are more informative than raw accepted length.
33. State and Memory Are Related but Distinct
State is the current compressed recurrent representation. Memory systems can additionally preserve external records, summaries or caches.
34. The State-Space Representation Audit
- What sequence history must remain available?
- What state size carries that history?
- How does current input change state dynamics?
- What can be forgotten safely?
- How are rare but important tokens protected?
- What long-range dependencies are tested?
- Does the model need random access to earlier detail?
- Can external retrieval restore forgotten evidence?
- Is the model used causally, bidirectionally or in streaming mode?
- How does state capacity affect compute and accuracy?
- Can state changes be replayed after correction?
- Does accepted sequence length correspond to usable retained information?
35. What Students Should Remember
- State-space models carry sequence history through recurrent latent state.
- Selective SSMs make state updates depend on input content.
- Mamba is a prominent selective state-space architecture.
- Recurrent state is compressed history, not a full archive.
- Long-sequence success depends on retaining the right information, not merely processing many tokens.
- State-space methods offer an efficient alternative to full attention and can be combined with retrieval or attention.
36. The Deep Principle
A state-space model remembers by transforming the past into a present state.
The art is not to carry everything forward. It is to preserve the parts of history that will still matter when the future finally asks for them.
