Glossary
Data

Data Lake

Also: Data Lake, Enterprise Data Lake

A data lake is a centralized repository that stores large volumes of raw data in its native format, from structured tables to unstructured files, until needed.

What It Is

A data lake is a centralized storage repository that holds vast amounts of data in its raw, native format. Unlike a traditional database or warehouse, it does not require you to define a schema before loading data. This is known as schema-on-read: structure is applied when the data is queried, not when it is stored.

A data lake can hold:

  • Structured data (relational tables, CSV files)
  • Semi-structured data (JSON, XML, log files, Parquet)
  • Unstructured data (images, video, audio, free text, PDFs)

Why it matters

Organizations generate data faster than they can model it. A data lake lets teams capture everything cheaply now and decide how to use it later. This flexibility supports advanced analytics and machine learning, which often need raw signals that a pre-aggregated warehouse would have discarded.

Key benefits:

  • Low cost storage using object stores (cloud or on-premises)
  • Decoupled storage and compute, so you scale each independently
  • Single source for analysts, data scientists, and reporting tools

The main risk is the data swamp: without governance, cataloging, and quality controls, a lake becomes an unsearchable dumping ground. Strong metadata, access policies, and a data catalog are essential.

How it is used in practice

Data lands in the lake through batch loads or streaming ingestion. It is commonly organized in zones: a raw zone (untouched source data), a cleansed zone (validated and deduplicated), and a curated zone (business ready datasets). Engines such as SQL query services, notebooks, and BI tools read directly from the lake.

Many teams now adopt the lakehouse pattern, which adds table formats and transactions on top of the lake to combine lake flexibility with warehouse reliability.

Concrete Example

A retailer streams clickstream logs, stores nightly sales tables, and dumps product images into one lake. Marketing queries clicks for campaign attribution, finance reads sales for reporting, and a data science team trains a recommendation model on clicks plus images, all from the same repository.

Data Lake ArchitectureDatabasesLogs / StreamsFiles / MediaSources (raw)Data LakeRaw zoneCleansedCuratedBI / ReportsML ModelsSQL QueriesConsumers
Raw data from many sources flows into zoned storage, then feeds analytics and ML consumers.

Frequently asked questions

What is the difference between a data lake and a data warehouse?

A data warehouse requires you to define the schema before loading data; a data lake stores data in its raw, native format and applies structure only when it is queried, an approach called schema-on-read. That makes the warehouse better suited to governed reporting on modeled data, and the lake better suited to keeping raw signals for exploration and machine learning. Cost profiles differ too: lakes rely on cheap object storage with storage and compute scaled separately.

What kinds of data can actually go into a data lake?

All three families. Structured data such as relational tables and CSV files, semi-structured data such as JSON, XML, log files and Parquet, and unstructured data such as images, video, audio, free text and PDFs. This is the point of the format: you capture everything cheaply now and decide how to model it later, instead of discarding signals a pre-aggregated warehouse would have dropped.

What turns a data lake into a data swamp, and how do you avoid it?

A lake becomes a data swamp when data accumulates without governance, cataloging or quality controls, leaving an unsearchable dumping ground nobody trusts. The safeguards are strong metadata, explicit access policies and a data catalog so users can find a dataset and know where it came from. Without those, low storage cost simply buys you a bigger problem.

How is data organized inside a data lake?

Most implementations use zones. A raw zone holds untouched source data, a cleansed zone holds validated and deduplicated data, and a curated zone holds business ready datasets. Data arrives through batch loads or streaming ingestion, and SQL query services, notebooks and BI tools read directly from these zones.

What does the lakehouse pattern add on top of a data lake?

The lakehouse adds table formats and transactions on top of the lake, so you keep the flexibility of raw storage while gaining the reliability normally associated with a warehouse. Many teams adopt it precisely to stop choosing between the two architectures. The underlying storage remains the same object store; what changes is the guarantees on reads and writes.