Three learners review open books together at a classroom table, with stacks of textbooks, stationery and a whiteboard in the bright room.

State-Space Representation | How Selective Recurrent State Carries Sequence History Without Full Attention

Attention remembers by keeping many earlier token representations available for comparison. State-space models take a different route: they carry sequence history forward through a recurrent internal state.

State-space representation is the problem of deciding what that state should preserve, what it should forget, how new input should update it and whether the resulting compressed history still contains the distinctions needed later.

This article continues the eduKateSingapore Representation and Tokenisation series beyond token creation, position and attention topology into recurrent sequence state.

The State-Space Route

TOKEN SEQUENCE
→ INPUT-DEPENDENT UPDATE
→ RECURRENT LATENT STATE
→ RETAIN / FORGET / TRANSFORM HISTORY
→ OUTPUT REPRESENTATION
→ NEXT TOKEN / CLASSIFICATION / CONTROL

STATE AT TIME t
= COMPRESSED HISTORY OF WHAT THE MODEL CHOSE TO CARRY FORWARD

1. Recurrent State Is a Compressed History

A recurrent sequence model does not necessarily keep every prior token representation explicitly available. Instead, it updates a state vector as new inputs arrive.

The state becomes a rolling summary of the past.

2. State Is Not the Same as Memory Archive

An archive preserves recoverable past records. A recurrent state preserves whatever transformed information the update rule carries forward.

Once information disappears from state, exact recovery may be impossible.

3. State-Space Models Come From Dynamical Systems

Classical state-space systems represent how a hidden state evolves over time under inputs and how outputs are generated from that state.

Modern neural SSMs adapt this idea to sequence modelling.

4. The Hidden State Is a Latent Representation

It need not correspond directly to words, events or human concepts.

Its job is operational: retain enough information about the sequence history to support future prediction or task decisions.

5. Linear State-Space Updates Can Be Highly Efficient

Unlike dense self-attention, which compares many token pairs, recurrent state updates can process a sequence with cost that scales linearly with sequence length.

This makes state-space methods attractive for long sequences.

6. Efficiency Alone Is Not Enough

A model can process a million tokens cheaply and still fail if the recurrent state forgets the one detail needed near the end.

Representation fidelity is more important than nominal sequence capacity.

7. Early SSMs Struggled With Content-Dependent Selection

Traditional linear state-space dynamics apply similar transition rules regardless of the specific token content.

Language and other discrete sequences often require selective memory: some tokens should strongly update state while others should be largely ignored.

8. Mamba Makes State Updates Input-Selective

Mamba: Linear-Time Sequence Modeling with Selective State Spaces makes key SSM parameters depend on the current input, allowing the model to decide what information to propagate or forget based on token content.

This turns state update into a learned content-sensitive filtering process.

9. Selectivity Is a Memory Gate

Not every token deserves equal influence on future state.

A selective state mechanism can amplify, suppress or transform incoming information before it enters long-term recurrent representation.

10. Selectivity Is Related to Adaptive Tokenisation

Adaptive tokenisation decides which representational units deserve continued compute. Selective state-space models decide which incoming information deserves persistent influence.

Both allocate representational life selectively.

11. Forgetting Is Built Into Recurrent State

Finite state cannot preserve every past distinction perfectly.

The update dynamics therefore implement an implicit forgetting policy even when no explicit memory-deletion operation exists.

12. Useful State Must Forget Redundancy Before Exceptions

Repeated formatting or predictable filler can be compressed aggressively.

A rare exception, negation or identity change may need persistent representation far into the future.

13. Long-Range Dependency Is a State-Retention Test

If a name introduced at token 500 determines the answer at token 50,000, the model must preserve the relevant identity through thousands of updates.

Long context is therefore a test of durable state, not just input acceptance.

14. Recurrent State Does Not Provide Random Access Automatically

Attention can directly revisit an earlier token representation if it remains in context.

A compressed recurrent state generally cannot “jump back” to exact earlier content unless that information was preserved in state or stored elsewhere.

15. This Changes the Meaning of Retrieval Inside the Sequence

An attention model can query explicit past token states. A state-space model queries its current summary of the past.

The distinction is direct memory versus compressed memory.

16. Compression Can Improve Robustness by Removing Distractors

A recurrent state that forgets irrelevant surface detail can focus on persistent structure.

The same compression can become harmful when the “surface detail” later turns out to matter.

17. State Dimension Is a Capacity Budget

A larger latent state can carry more information and costs more memory and compute.

A smaller state forces stronger compression and greater competition among remembered features.

18. State Capacity Should Match Task Complexity

Simple local pattern recognition may need little state. Long-range program analysis, genomics or document reasoning may require substantially richer state.

There is no universal state size independent of receiver job.

19. State Update Frequency Matters

Updating at every primitive byte, token or frame gives fine temporal resolution.

Updating over chunks or pooled representations reduces sequence length while shifting detail into the chunk encoder.

20. Hierarchical State Can Operate at Several Timescales

Fast state can capture local syntax or motion; slower state can track topic, scene or process stage.

Multi-timescale recurrence mirrors Hierarchical Representation.

21. State-Space Models Can Process Multiple Modalities

Mamba reports strong sequence-model performance across language, audio and genomics.

The common abstraction is a sequence whose useful history can be carried through selective state.

22. Modality Still Determines What State Must Preserve

Audio needs temporal phase and phonetic structure; genomics needs motif and long-range regulatory context; language needs entities, syntax and discourse.

One architecture can span modalities without making their representational jobs identical.

23. State-Space Models Change Interaction Topology

Information does not move through arbitrary token-to-token attention edges.

It moves forward through state transitions, creating a sequential communication topology.

24. This Is a Different Path Length

Two distant tokens can influence each other only through the chain of intervening state updates.

Selective dynamics must preserve the relevant signal over that path.

25. Bidirectional Use Requires Two Directions or Another Design

A left-to-right recurrent state naturally uses past context. Whole-sequence representation tasks can process both directions or combine recurrent passes.

Directionality should match the information available at inference time.

26. Streaming Is a Natural Strength

Because state can be updated incrementally, a state-space model can process streaming audio, sensor data or long text without storing an ever-growing explicit history.

The representation is naturally online.

27. Streaming Makes Correction Harder

If an earlier input is later discovered to be wrong, a recurrent state may need replay from that point or an explicit correction mechanism.

Compressed history is efficient and less editable.

28. External Retrieval Complements Recurrent State

Facts that are too important to risk forgetting can live in an external store and be retrieved when needed.

State carries working history; retrieval restores addressable source detail.

29. Hybrid Architectures Can Combine Attention and State Space

Attention is strong at direct content-based interaction; recurrent state is strong at efficient sequential compression.

Architectures can mix these mechanisms rather than treat them as mutually exclusive.

30. State Interpretability Is Difficult

A recurrent vector does not come with labels saying which dimensions represent topic, identity or timing.

Probing can reveal information encoded in state without proving the model uses that information causally.

31. State Failure Is Often Invisible Until a Later Query

The model may appear coherent for thousands of steps after silently forgetting a critical premise.

Evaluation should place decisive information far in the past and test whether it remains usable.

32. Sequence Length Benchmarks Should Test Content Retention

Processing a million positions is not equivalent to remembering useful information across a million positions.

Needle retrieval, multi-step dependencies and long-range state changes are more informative than raw accepted length.

33. State and Memory Are Related but Distinct

State is the current compressed recurrent representation. Memory systems can additionally preserve external records, summaries or caches.

See Memory Representation.

34. The State-Space Representation Audit

  1. What sequence history must remain available?
  2. What state size carries that history?
  3. How does current input change state dynamics?
  4. What can be forgotten safely?
  5. How are rare but important tokens protected?
  6. What long-range dependencies are tested?
  7. Does the model need random access to earlier detail?
  8. Can external retrieval restore forgotten evidence?
  9. Is the model used causally, bidirectionally or in streaming mode?
  10. How does state capacity affect compute and accuracy?
  11. Can state changes be replayed after correction?
  12. Does accepted sequence length correspond to usable retained information?

35. What Students Should Remember

36. The Deep Principle

A state-space model remembers by transforming the past into a present state.

The art is not to carry everything forward. It is to preserve the parts of history that will still matter when the future finally asks for them.

Continue the Representation & Tokenisation Series

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate SG

Subscribe now to keep reading and get access to the full archive.

Continue reading