Data Lakes and Lakehouses
A data lake is an analytical storage architecture designed to hold large volumes of structured, semi-structured and unstructured data in relatively flexible forms, commonly on scalable object storage. A lakehouse adds stronger table, transaction, metadata and governance capabilities so that lake-style storage can support more warehouse-like reliability and analytical use.
A lake is useful when it preserves optionality. A lakehouse becomes useful when that optionality is disciplined enough that people can trust what they query.
The attraction is straightforward: organisations want one large analytical estate where raw files, curated datasets, machine-learning features, logs, documents and historical data can coexist without forcing every object into one rigid database design at the moment of arrival. The danger is equally straightforward: flexible storage without strong metadata, ownership, quality and lifecycle control becomes a data swamp.
ARTICLE ID: DATA.MANAGEMENT.029
Canonical function: large-scale analytical storage, layered refinement and governed lakehouse architecture
Series route: Data Architecture → Data Lakes and Lakehouses.
The Simple Answer
A data lake stores many kinds of data cheaply and flexibly. A lakehouse adds enough structure that those files can behave more like dependable analytical tables.
A useful conceptual route is:
Source → Land → Preserve → Validate → Curate → Model → Publish → Observe → Retain
The lake is the storage landscape. The lakehouse is the combination of storage, table abstractions, metadata, governance and processing that makes the landscape reliably usable.
Why Object Storage Matters
Object storage is well suited to large analytical estates because it separates durable storage from individual compute engines. Files can remain in one underlying storage layer while different processing engines read and transform them.
This separation creates several advantages:
- storage can scale independently from compute;
- historical data can be retained economically;
- several tools can work over the same underlying objects;
- raw source files can be preserved before transformation;
- compute can be increased temporarily for heavy jobs and reduced afterwards.
But low-cost storage can encourage careless accumulation. Cheap storage is not the same as cheap stewardship.
Files Are Not Tables
Object storage holds objects: files and metadata about those files. Analytical users usually want tables with stable columns, partitions, versions and predictable update behaviour.
A lakehouse table abstraction sits above the files and records which files belong to a table version, which schema applies and which changes have been committed. This can provide transaction-like behaviour over object storage without turning the object store itself into a traditional database.
Open Table Abstractions
Modern lakehouse designs commonly use open or interoperable table abstractions that describe table state, schema evolution, snapshots and partition information separately from individual processing engines.
The architectural benefit is reduced lock-in at the data layer: several engines may be able to understand the same underlying table representation. The practical benefit depends on actual compatibility, operations and governance; “open” does not mean every tool behaves identically.
Transaction-Like Guarantees
Without a table transaction layer, a multi-file update can expose partial state. One query may see some new files and some old ones. A lakehouse table can publish a new snapshot atomically at the metadata level so readers see either the previous table state or the newly committed state.
This matters because analytical correctness depends on coherent versions, not merely durable files.
The Raw Zone
A raw or landing zone preserves source data close to the form in which it arrived. Its purpose is evidence and recoverability, not immediate convenience.
- preserve source identity;
- record ingestion time;
- retain source timestamps;
- capture file manifests or checksums where useful;
- avoid silently rewriting values;
- apply appropriate security and retention from the moment of arrival.
“Raw” should not be interpreted as ungoverned. Sensitive raw data can be the highest-risk part of the estate.
Bronze, Silver and Gold
Many teams use a bronze–silver–gold metaphor for progressive refinement:
- Bronze: landed or minimally transformed source data.
- Silver: validated, standardised, reconciled and conformed data.
- Gold: business-ready or analytical products designed for specific consumers.
The names are not important. The principle is: preserve source evidence, make quality changes explicit, and separate intermediate preparation from trusted receiver-facing products.
Do Not Confuse Layers with Trust Automatically
A dataset does not become trustworthy because it lives in a folder called gold. Trust should come from ownership, contract, lineage, validation, freshness and actual receiver fitness.
Layer names are architecture. Trust is an evidence state.
Schema-on-Read and Schema-on-Write
Schema-on-write validates and structures data before or during admission to a governed table. Schema-on-read applies structure when data is queried or interpreted.
Lakes are often associated with schema-on-read, while warehouses are associated with schema-on-write. In practice, mature lakehouses use both. Raw zones may preserve flexible source forms while curated tables enforce stronger schemas.
Schema Evolution
Analytical datasets change. Columns are added, types evolve, nested structures change and definitions are revised.
A safe lakehouse design distinguishes:
- additive changes;
- compatible changes;
- breaking changes;
- semantic changes that leave the physical schema unchanged.
See Data Versioning and Change Management.
Partitioning
Partitioning groups data according to values such as date, region or domain so engines can avoid scanning irrelevant files.
Good partitioning reduces read cost. Poor partitioning creates excessive small partitions or fails to match actual query patterns.
The Small-Files Problem
Streaming and frequent micro-batches can create enormous numbers of tiny files. Even when total storage volume is modest, metadata operations and query planning can become expensive.
Compaction combines small files into more efficient larger files while preserving table state. Compaction is an operational maintenance task, not a one-time setup decision.
File Size Is a Trade-Off
Very small files increase overhead. Very large files can reduce parallelism or make selective reads expensive. The useful size depends on engine behaviour, query patterns, update frequency and partitioning.
There is no universal ideal file size independent of workload.
Columnar Storage
Analytical lakehouses often use columnar file formats because analytical queries frequently read a subset of columns over many rows. Columnar layout can reduce I/O and improve compression for such workloads.
The file format is one layer. Correctness still depends on table metadata, definitions, quality and lineage.
Metadata Is the Lake’s Navigation System
Without metadata, object storage is a large collection of paths. A usable lakehouse needs metadata describing datasets, schemas, owners, partitions, versions, classifications, quality, lineage and lifecycle.
See Metadata and Data Lineage.
The Catalogue
A catalogue gives people and machines a searchable inventory of the estate. It should help distinguish raw, curated, trusted, deprecated and sensitive datasets.
See Data Catalogues and Discovery.
Canonical Sources
A lakehouse can accidentally create several versions of the same concept: raw customer records, cleansed customer tables, analytics customer dimensions and machine-learning customer features.
Architecture should explain which dataset owns identity, which owns analytical history, and which are derived products. A large lake does not remove the need for canonical ownership.
Warehouse vs Lake vs Lakehouse
A traditional data warehouse emphasises structured analytical modelling, controlled schemas and business-ready data. A lake emphasises flexible, scalable storage of diverse data. A lakehouse tries to combine flexible storage with stronger table reliability, transactions, metadata and governance.
These are architectural patterns, not moral rankings. An organisation may use warehouses, lakes and lakehouses together.
See Data Warehousing and Analytics.
Data Marts
Even inside a lakehouse, consumer-specific marts or semantic models can remain useful. A central lake does not mean every analyst should query raw tables directly.
Good architecture separates durable shared foundations from receiver-friendly interfaces.
Batch and Streaming Together
Lakehouses often ingest both batch files and event streams. A shared table layer can make streaming updates visible to analytical queries while still supporting historical recomputation.
The difficulty is maintaining consistent event time, deduplication, late-arriving data and version semantics across both modes.
Late-Arriving Data
Historical partitions can change after their first publication because events arrive late, source systems correct records or backfills occur.
Consumers should know whether a daily partition is provisional, final or still subject to restatement.
Time Travel
Snapshot-based tables can allow queries against prior table states. This is often called time travel.
Time travel can support debugging, reproducibility and rollback, but it is not infinite archival preservation by default. Old snapshots and files may eventually be expired under retention rules.
Retention and Vacuuming
Table snapshots can keep old files reachable for a period. Maintenance may eventually remove files no longer required by active snapshots.
Retention must balance recovery, reproducibility, cost, privacy and legal obligations. Deleting old physical files too aggressively can destroy rollback or reproducibility; retaining them forever can create cost and exposure.
Governance Must Start at Landing
Waiting until data reaches a curated zone to apply governance is too late. Raw files can already contain personal, confidential or restricted information.
Classification, access, lineage and retention should begin at admission.
See Data Classification and Sensitivity.
Fine-Grained Access
Lakehouses may need controls at dataset, table, row, column or view level. The correct granularity depends on data sensitivity and receiver need.
Derived tables can sometimes reduce exposure by publishing only the fields required for a particular analytical purpose.
Encryption
Encryption at rest and in transit protects confidentiality, but the lakehouse also needs strong identity, key management and access governance. Encryption does not stop an overprivileged authorised user from reading data.
Data Sovereignty
Object stores, compute engines, catalogues and backups can live in different regions or services. Sovereignty review should map the entire lakehouse control and data plane.
See Data Sovereignty, Residency and Jurisdiction.
Data Quality in the Lakehouse
Quality gates can operate between layers:
- schema checks at landing;
- deduplication before curation;
- reference-data validation;
- business invariants;
- source-to-target reconciliation;
- freshness checks;
- consumer contract tests.
Bad rows should not simply disappear during refinement. Quarantine and exception metadata preserve the evidence needed for repair.
Observability
A healthy lakehouse monitors more than compute jobs. It observes table freshness, partition arrival, volume, schema, quality, file growth, compaction health and lineage impact.
See Data Observability and Monitoring.
Cost Governance
Separating compute from storage improves flexibility but can make cost harder to see. A cheap object store can support expensive scans if queries repeatedly read huge volumes.
Cost controls may include:
- partition pruning;
- column pruning;
- compaction;
- appropriate file formats;
- workload limits;
- materialised or precomputed products for repeated queries;
- retention rules for obsolete intermediate data;
- chargeback or cost attribution by domain.
Data Swamps
A data lake becomes a swamp when people cannot tell what exists, which datasets are current, who owns them, what they mean or whether they are safe to use.
The antidote is not more folders. It is ownership, metadata, lineage, lifecycle, contracts and quality.
Shadow Lakes
Teams can create private object-storage areas or notebook outputs that become unofficial analytical estates. Shadow lakes recreate the same governance problems that central lakes were meant to solve.
Useful experimentation should be supported, but experimental data should have visible lifecycle states and promotion routes into governed products.
AI and Machine Learning
Lakes and lakehouses are attractive for AI because they can hold documents, images, logs, labels, structured features and historical observations in one broad analytical estate.
AI increases the need for:
- training-data provenance;
- versioned datasets;
- rights and licence metadata;
- label quality;
- population coverage;
- separation of training and evaluation data;
- controlled access to sensitive corpora;
- reproducible feature generation.
Vector Data and Lakehouses
Embeddings and vector indexes may be derived from documents stored in the lakehouse. The important management principle is to preserve a route from each derived vector representation back to source document identity, version and access rights.
A vector index should not become an orphaned shadow copy of organisational knowledge.
Education Example
An education lakehouse may land attendance, assessment, curriculum, learning-platform and communication data. Raw source records are preserved. Curated layers standardise student and subject identities. Trusted analytical products expose only authorised fields needed for teaching and planning.
The architecture succeeds when analysts do not need to reconstruct student identity and definitions independently for every report.
Research Example
A research lakehouse can preserve instrument files, metadata, derived tables and computational outputs. Snapshot versioning can make analyses reproducible while raw data remains available for reprocessing under improved methods.
Commercial Example
A retailer can combine transactions, inventory, web events, returns and product information inside one governed estate. Raw events support exploration; conformed product and customer identities support reliable analytics; gold products support forecasting and finance.
Common Failure Modes
- Store everything, govern later: raw storage becomes unmanaged risk.
- Folder equals architecture: naming conventions substitute for metadata and contracts.
- Gold by label: trusted status is asserted rather than evidenced.
- Schema freedom forever: every consumer interprets source files differently.
- Small-file explosion: operational overhead grows faster than data volume.
- No compaction: query efficiency deteriorates over time.
- No canonical ownership: several curated versions of the same entity compete.
- Lakehouse equals warehouse replacement: useful marts and semantic layers are discarded unnecessarily.
- Infinite time travel: old snapshots accumulate without retention policy.
- AI shadow copies: embeddings and model corpora lose source identity and permissions.
A Lakehouse Design Checklist
- Which data belongs in the lakehouse?
- What should be preserved in raw form?
- Which table abstraction and metadata model are used?
- How are schemas and versions managed?
- How are partitions designed?
- How is small-file growth controlled?
- Which datasets are canonical?
- How are quality gates enforced between layers?
- How are sensitive datasets classified from landing onward?
- How are catalogue and lineage maintained?
- What snapshot retention is required?
- How are batch and streaming changes reconciled?
- How are compute and storage costs attributed?
- How are experimental datasets promoted or retired?
- Can a receiver trace a trusted analytical result back to source evidence?
A Maturity Ladder
- Stored: diverse files are centralised.
- Layered: raw, curated and consumer-ready states are separated.
- Tabular: governed table abstractions provide stable snapshots.
- Catalogued: ownership, metadata and discovery are operational.
- Quality-gated: promotion between layers requires evidence.
- Observable: freshness, files, partitions and quality are monitored.
- Product-oriented: curated outputs have contracts and receivers.
- Adaptive: usage, incidents and cost continuously improve the estate.
The Deeper Principle: Preserve Optionality Without Preserving Chaos
The fundamental promise of a data lake is optionality: keep enough source evidence that future questions can still be asked. The fundamental promise of a lakehouse is controlled optionality: preserve that breadth without forcing every receiver to rediscover meaning, quality and history from raw files.
The winning architecture is therefore not the one with the largest lake. It is the one where raw evidence, curated meaning and trusted products remain visibly connected.
Data Management Series
- Data Lakes and Lakehouses
- Data Architecture
- Data Warehousing and Analytics
- Data Engineering and Pipelines
- Data Streaming and Event-Driven Systems
Final idea: a lakehouse is not a giant bucket with a fashionable name. It is a governed analytical memory where object storage, table state, metadata, quality, security and lifecycle work together so flexible storage becomes dependable knowledge infrastructure.
