Data Lakes and Lakehouses | Object Storage, Open Tables, Governance and Analytical Scale

Data Lakes and Lakehouses

A data lake is an analytical storage architecture designed to hold large volumes of structured, semi-structured and unstructured data in relatively flexible forms, commonly on scalable object storage. A lakehouse adds stronger table, transaction, metadata and governance capabilities so that lake-style storage can support more warehouse-like reliability and analytical use.

A lake is useful when it preserves optionality. A lakehouse becomes useful when that optionality is disciplined enough that people can trust what they query.

The attraction is straightforward: organisations want one large analytical estate where raw files, curated datasets, machine-learning features, logs, documents and historical data can coexist without forcing every object into one rigid database design at the moment of arrival. The danger is equally straightforward: flexible storage without strong metadata, ownership, quality and lifecycle control becomes a data swamp.

ARTICLE ID: DATA.MANAGEMENT.029
Canonical function: large-scale analytical storage, layered refinement and governed lakehouse architecture
Series route: Data Architecture → Data Lakes and Lakehouses.

The Simple Answer

A data lake stores many kinds of data cheaply and flexibly. A lakehouse adds enough structure that those files can behave more like dependable analytical tables.

A useful conceptual route is:

Source → Land → Preserve → Validate → Curate → Model → Publish → Observe → Retain

The lake is the storage landscape. The lakehouse is the combination of storage, table abstractions, metadata, governance and processing that makes the landscape reliably usable.

Why Object Storage Matters

Object storage is well suited to large analytical estates because it separates durable storage from individual compute engines. Files can remain in one underlying storage layer while different processing engines read and transform them.

This separation creates several advantages:

But low-cost storage can encourage careless accumulation. Cheap storage is not the same as cheap stewardship.

Files Are Not Tables

Object storage holds objects: files and metadata about those files. Analytical users usually want tables with stable columns, partitions, versions and predictable update behaviour.

A lakehouse table abstraction sits above the files and records which files belong to a table version, which schema applies and which changes have been committed. This can provide transaction-like behaviour over object storage without turning the object store itself into a traditional database.

Open Table Abstractions

Modern lakehouse designs commonly use open or interoperable table abstractions that describe table state, schema evolution, snapshots and partition information separately from individual processing engines.

The architectural benefit is reduced lock-in at the data layer: several engines may be able to understand the same underlying table representation. The practical benefit depends on actual compatibility, operations and governance; “open” does not mean every tool behaves identically.

Transaction-Like Guarantees

Without a table transaction layer, a multi-file update can expose partial state. One query may see some new files and some old ones. A lakehouse table can publish a new snapshot atomically at the metadata level so readers see either the previous table state or the newly committed state.

This matters because analytical correctness depends on coherent versions, not merely durable files.

The Raw Zone

A raw or landing zone preserves source data close to the form in which it arrived. Its purpose is evidence and recoverability, not immediate convenience.

“Raw” should not be interpreted as ungoverned. Sensitive raw data can be the highest-risk part of the estate.

Bronze, Silver and Gold

Many teams use a bronze–silver–gold metaphor for progressive refinement:

The names are not important. The principle is: preserve source evidence, make quality changes explicit, and separate intermediate preparation from trusted receiver-facing products.

Do Not Confuse Layers with Trust Automatically

A dataset does not become trustworthy because it lives in a folder called gold. Trust should come from ownership, contract, lineage, validation, freshness and actual receiver fitness.

Layer names are architecture. Trust is an evidence state.

Schema-on-Read and Schema-on-Write

Schema-on-write validates and structures data before or during admission to a governed table. Schema-on-read applies structure when data is queried or interpreted.

Lakes are often associated with schema-on-read, while warehouses are associated with schema-on-write. In practice, mature lakehouses use both. Raw zones may preserve flexible source forms while curated tables enforce stronger schemas.

Schema Evolution

Analytical datasets change. Columns are added, types evolve, nested structures change and definitions are revised.

A safe lakehouse design distinguishes:

See Data Versioning and Change Management.

Partitioning

Partitioning groups data according to values such as date, region or domain so engines can avoid scanning irrelevant files.

Good partitioning reduces read cost. Poor partitioning creates excessive small partitions or fails to match actual query patterns.

The Small-Files Problem

Streaming and frequent micro-batches can create enormous numbers of tiny files. Even when total storage volume is modest, metadata operations and query planning can become expensive.

Compaction combines small files into more efficient larger files while preserving table state. Compaction is an operational maintenance task, not a one-time setup decision.

File Size Is a Trade-Off

Very small files increase overhead. Very large files can reduce parallelism or make selective reads expensive. The useful size depends on engine behaviour, query patterns, update frequency and partitioning.

There is no universal ideal file size independent of workload.

Columnar Storage

Analytical lakehouses often use columnar file formats because analytical queries frequently read a subset of columns over many rows. Columnar layout can reduce I/O and improve compression for such workloads.

The file format is one layer. Correctness still depends on table metadata, definitions, quality and lineage.

Metadata Is the Lake’s Navigation System

Without metadata, object storage is a large collection of paths. A usable lakehouse needs metadata describing datasets, schemas, owners, partitions, versions, classifications, quality, lineage and lifecycle.

See Metadata and Data Lineage.

The Catalogue

A catalogue gives people and machines a searchable inventory of the estate. It should help distinguish raw, curated, trusted, deprecated and sensitive datasets.

See Data Catalogues and Discovery.

Canonical Sources

A lakehouse can accidentally create several versions of the same concept: raw customer records, cleansed customer tables, analytics customer dimensions and machine-learning customer features.

Architecture should explain which dataset owns identity, which owns analytical history, and which are derived products. A large lake does not remove the need for canonical ownership.

Warehouse vs Lake vs Lakehouse

A traditional data warehouse emphasises structured analytical modelling, controlled schemas and business-ready data. A lake emphasises flexible, scalable storage of diverse data. A lakehouse tries to combine flexible storage with stronger table reliability, transactions, metadata and governance.

These are architectural patterns, not moral rankings. An organisation may use warehouses, lakes and lakehouses together.

See Data Warehousing and Analytics.

Data Marts

Even inside a lakehouse, consumer-specific marts or semantic models can remain useful. A central lake does not mean every analyst should query raw tables directly.

Good architecture separates durable shared foundations from receiver-friendly interfaces.

Batch and Streaming Together

Lakehouses often ingest both batch files and event streams. A shared table layer can make streaming updates visible to analytical queries while still supporting historical recomputation.

The difficulty is maintaining consistent event time, deduplication, late-arriving data and version semantics across both modes.

Late-Arriving Data

Historical partitions can change after their first publication because events arrive late, source systems correct records or backfills occur.

Consumers should know whether a daily partition is provisional, final or still subject to restatement.

Time Travel

Snapshot-based tables can allow queries against prior table states. This is often called time travel.

Time travel can support debugging, reproducibility and rollback, but it is not infinite archival preservation by default. Old snapshots and files may eventually be expired under retention rules.

Retention and Vacuuming

Table snapshots can keep old files reachable for a period. Maintenance may eventually remove files no longer required by active snapshots.

Retention must balance recovery, reproducibility, cost, privacy and legal obligations. Deleting old physical files too aggressively can destroy rollback or reproducibility; retaining them forever can create cost and exposure.

Governance Must Start at Landing

Waiting until data reaches a curated zone to apply governance is too late. Raw files can already contain personal, confidential or restricted information.

Classification, access, lineage and retention should begin at admission.

See Data Classification and Sensitivity.

Fine-Grained Access

Lakehouses may need controls at dataset, table, row, column or view level. The correct granularity depends on data sensitivity and receiver need.

Derived tables can sometimes reduce exposure by publishing only the fields required for a particular analytical purpose.

Encryption

Encryption at rest and in transit protects confidentiality, but the lakehouse also needs strong identity, key management and access governance. Encryption does not stop an overprivileged authorised user from reading data.

Data Sovereignty

Object stores, compute engines, catalogues and backups can live in different regions or services. Sovereignty review should map the entire lakehouse control and data plane.

See Data Sovereignty, Residency and Jurisdiction.

Data Quality in the Lakehouse

Quality gates can operate between layers:

Bad rows should not simply disappear during refinement. Quarantine and exception metadata preserve the evidence needed for repair.

Observability

A healthy lakehouse monitors more than compute jobs. It observes table freshness, partition arrival, volume, schema, quality, file growth, compaction health and lineage impact.

See Data Observability and Monitoring.

Cost Governance

Separating compute from storage improves flexibility but can make cost harder to see. A cheap object store can support expensive scans if queries repeatedly read huge volumes.

Cost controls may include:

Data Swamps

A data lake becomes a swamp when people cannot tell what exists, which datasets are current, who owns them, what they mean or whether they are safe to use.

The antidote is not more folders. It is ownership, metadata, lineage, lifecycle, contracts and quality.

Shadow Lakes

Teams can create private object-storage areas or notebook outputs that become unofficial analytical estates. Shadow lakes recreate the same governance problems that central lakes were meant to solve.

Useful experimentation should be supported, but experimental data should have visible lifecycle states and promotion routes into governed products.

AI and Machine Learning

Lakes and lakehouses are attractive for AI because they can hold documents, images, logs, labels, structured features and historical observations in one broad analytical estate.

AI increases the need for:

Vector Data and Lakehouses

Embeddings and vector indexes may be derived from documents stored in the lakehouse. The important management principle is to preserve a route from each derived vector representation back to source document identity, version and access rights.

A vector index should not become an orphaned shadow copy of organisational knowledge.

Education Example

An education lakehouse may land attendance, assessment, curriculum, learning-platform and communication data. Raw source records are preserved. Curated layers standardise student and subject identities. Trusted analytical products expose only authorised fields needed for teaching and planning.

The architecture succeeds when analysts do not need to reconstruct student identity and definitions independently for every report.

Research Example

A research lakehouse can preserve instrument files, metadata, derived tables and computational outputs. Snapshot versioning can make analyses reproducible while raw data remains available for reprocessing under improved methods.

Commercial Example

A retailer can combine transactions, inventory, web events, returns and product information inside one governed estate. Raw events support exploration; conformed product and customer identities support reliable analytics; gold products support forecasting and finance.

Common Failure Modes

A Lakehouse Design Checklist

  1. Which data belongs in the lakehouse?
  2. What should be preserved in raw form?
  3. Which table abstraction and metadata model are used?
  4. How are schemas and versions managed?
  5. How are partitions designed?
  6. How is small-file growth controlled?
  7. Which datasets are canonical?
  8. How are quality gates enforced between layers?
  9. How are sensitive datasets classified from landing onward?
  10. How are catalogue and lineage maintained?
  11. What snapshot retention is required?
  12. How are batch and streaming changes reconciled?
  13. How are compute and storage costs attributed?
  14. How are experimental datasets promoted or retired?
  15. Can a receiver trace a trusted analytical result back to source evidence?

A Maturity Ladder

  1. Stored: diverse files are centralised.
  2. Layered: raw, curated and consumer-ready states are separated.
  3. Tabular: governed table abstractions provide stable snapshots.
  4. Catalogued: ownership, metadata and discovery are operational.
  5. Quality-gated: promotion between layers requires evidence.
  6. Observable: freshness, files, partitions and quality are monitored.
  7. Product-oriented: curated outputs have contracts and receivers.
  8. Adaptive: usage, incidents and cost continuously improve the estate.

The Deeper Principle: Preserve Optionality Without Preserving Chaos

The fundamental promise of a data lake is optionality: keep enough source evidence that future questions can still be asked. The fundamental promise of a lakehouse is controlled optionality: preserve that breadth without forcing every receiver to rediscover meaning, quality and history from raw files.

The winning architecture is therefore not the one with the largest lake. It is the one where raw evidence, curated meaning and trusted products remain visibly connected.

Data Management Series


Final idea: a lakehouse is not a giant bucket with a fashionable name. It is a governed analytical memory where object storage, table state, metadata, quality, security and lifecycle work together so flexible storage becomes dependable knowledge infrastructure.

Explore the connected learning guides

Choose the question that brought you here. Open one useful guide, try a small task, and stop when you have what you need.

Take one question further

The same learning habit can travel across subjects, while each subject keeps its own methods. These routes help you notice a difficulty, understand one part of it, and return to something you can do.

A word is familiar, but using it is difficult.

Move from recognising a word to retrieving it in a new context. Understand vocabulary plateaus.

Try it without the guide: Choose one word you already know. Close the guide and use it in a new sentence. Explain why it fits; try another context tomorrow.

A piece of writing has ideas, but the reader loses the thread.

Make the order of events and the links between sentences clear. Explore composition writing.

Try it without the guide: Choose one short paragraph. Read the relevant explanation, close it, and revise the paragraph. Ask someone to tell you what happened and why.

The Mathematics seems familiar, but marks still disappear.

Find the first point where the working stops being reliable. Find Secondary 4 A-Math mark leakage.

Try it without the guide: For a Secondary 4 A-Math question you have attempted, locate the first uncertain line. Repair that step, then try a comparable question without the worked answer.

A Science fact is remembered, but the explanation is incomplete.

Connect the evidence to a scientific idea and the resulting change. Follow the Primary Science learning route.

Try it without the guide: Choose a familiar Primary Science example. Explain the evidence, the idea and the result without notes. Then change one condition and explain your prediction.

Two accounts of the world seem to disagree.

Check the question, source, date and evidence before combining claims. Explore the World Knowledge research library.

Try it without the guide: Take one claim. Find the source best placed to support it, note its date, and state what remains uncertain. Return to your original question.

There is plenty of help, but independence is hard to see.

Check what the learner can understand and do after support is removed. Understand how education works.

Try it without the guide: Choose one small task the child has practised. Agree on a calm, brief attempt without prompts. Use what happens to choose one next step, then stop.

For the structure behind these connections, read the eduKateSingapore runtime manifest and the eduKate ecosystem boot contract. The reader map describes public navigation; those manifests preserve the wider ownership and return rules.

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading