Glossary
DataAIgeneral

Data pipeline

Also: data pipeline, pipeline, ETL pipeline, ELT pipeline, data flow, pipeline de données, chaîne de traitement des données

An automated sequence of steps that moves data from source to destination: ingestion, transformation, validation, and loading, so it arrives clean and ready to use.

What it is

A data pipeline is an automated series of steps that carries data from where it originates (a source) to where it is consumed (a destination), applying processing along the way. The core stages are usually:

  • Ingestion: pulling or receiving raw data from sources such as databases, APIs, files, event streams, or SaaS tools.
  • Transformation: cleaning, reshaping, joining, and enriching the data into a usable form.
  • Validation: checking quality, completeness, and conformity to expected rules before the data moves on.
  • Loading: writing the result into a destination such as a data warehouse, lake, dashboard, or application.

Pipelines can run in batch (scheduled, for example hourly or nightly) or in streaming mode (continuously, near real time).

Why it matters

Without pipelines, data stays trapped in silos and every analysis becomes a manual, error-prone copy-paste exercise. A well-built pipeline delivers:

  • Reliability: the same steps run the same way every time, with monitoring and alerts.
  • Freshness: decisions rest on current data, not last quarter's export.
  • Trust: validation catches bad data before it reaches reports or models.
  • Scale: volumes that a person could never process by hand flow automatically.

For executives, the pipeline is the plumbing behind every dashboard, forecast, and AI feature. When it breaks silently, numbers drift and confidence erodes.

How it is used in practice

Teams describe pipelines as code (often called orchestration), schedule them, and monitor each run. Key practical concerns include handling failures gracefully, reprocessing when a source changes, tracking lineage (where each number came from), and controlling cost.

Worked example

A retailer wants a daily revenue-by-region dashboard:

1. Ingestion: each night the pipeline pulls orders from the e-commerce database and exchange rates from a currency API.

2. Transformation: it converts every order to euros, joins each order to its store region, and aggregates totals.

3. Validation: it checks that no region is missing and that today's total is within a plausible range of yesterday's. If a rule fails, it alerts the data team and stops.

4. Loading: clean, aggregated figures land in the warehouse, and the dashboard refreshes.

The CFO sees trustworthy numbers by 8 a.m., without anyone touching a spreadsheet.

SourcesDatabaseAPIFilesIngestionTransformationValidationLoadingWarehouse / Dashboard
Data flows from multiple sources through ingestion, transformation, and validation, then loads into a warehouse or dashboard.

Frequently asked questions

What is a data pipeline?

A data pipeline is an automated sequence of steps that carries data from a source (database, API, file, event stream, SaaS tool) to a destination such as a data warehouse, a dashboard, or an application. Along the way it applies four stages: ingestion, transformation, validation, and loading. The point is that the data arrives clean and ready to use, without anyone copying and pasting.

Why should a non-technical executive care about pipelines?

Because the pipeline is the plumbing behind every dashboard, forecast, and AI feature you rely on. When it breaks silently, numbers drift and trust in reporting erodes, usually before anyone notices. Understanding the four stages lets you ask the right question when a figure looks wrong: was it ingestion, transformation, validation, or loading?

What is the difference between batch and streaming pipelines?

A batch pipeline runs on a schedule, for example hourly or nightly, and processes accumulated data in one pass. A streaming pipeline runs continuously and delivers data in near real time. Batch suits daily reporting like a revenue dashboard; streaming suits cases where a delay of hours would make the data useless.

What happens if a validation rule fails during a pipeline run?

A well-built pipeline stops and alerts the data team rather than loading suspect figures. In a daily revenue dashboard, for instance, validation checks that no region is missing and that today's total stays within a plausible range of yesterday's; if a rule fails, the run halts before the dashboard refreshes. That is precisely what protects reports from silently wrong numbers.

What does data lineage mean and why is it tracked in pipelines?

Lineage is the traceability of each number: which source it came from and which transformations it went through. Teams track it so that when a figure is questioned or a source schema changes, they can identify what to reprocess and what else is affected. It sits alongside failure handling and cost control among the practical concerns of running pipelines in production.