+55 XP

Pipelines de données : ETL/ELT, batch, streaming et architecture Medallion

Data pipelines are the circulatory system of your data platform. They move data from where it's produced to where it's consumed. When they work, nobody notices. When they fail, everything stops.

A CDO who doesn't understand data pipeline architecture can't diagnose failures, can't have credible conversations with engineering teams, and can't make good architectural decisions.

ETL vs. ELT: the fundamental pattern shift

Traditional ETL (Extract, Transform, Load) transforms data before loading it to the destination. ELT (Extract, Load, Transform) loads raw data first, then transforms it inside the destination system.

ETL made sense when compute was expensive and storage was cheap, you minimized what you stored. ELT makes sense when cloud warehouses (Snowflake, BigQuery) offer scalable compute, and you want to preserve raw data for flexible reuse.

Modern default: ELT. Load raw data to your warehouse or lake. Transform using SQL (dbt) or Spark. Re-transform as needs evolve without reloading source data.

ETL vs ELT: Key Differences Explained

Watch on YouTube

Knowledge check

1. What is the fundamental distinction between ETL and ELT?

2. Why has ELT become the modern default over traditional ETL?

3. A fraud detection system must react to transactions as they happen. Which processing pattern is most appropriate?

MULTIPLE CHOICE

4. Select ALL correct statements about batch, streaming, and micro-batch processing.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL correct statements about pipeline orchestration and Apache Airflow.

Select all the correct answers.

Batch vs. streaming: the latency spectrum

Batch processing moves data in scheduled windows, hourly, daily, weekly. Tools: Apache Spark, dbt, AWS Glue. Use when: latency is acceptable (overnight reports, weekly models), data volume is high, and processing is complex.

Streaming processing moves data in real-time or near-real-time (seconds to minutes). Tools: Apache Kafka, Apache Flink, AWS Kinesis, Google Pub/Sub. Use when: business decisions depend on current data (fraud detection, real-time personalization, operational dashboards).

Micro-batch is the middle ground, small batches processed every few minutes. Apache Spark Structured Streaming supports this. Often sufficient when true real-time isn't required but hourly batch is too slow.

Most organizations need both: batch for heavy analytical workloads, streaming for operational use cases. Architect to support both patterns, don't build separate, incompatible systems for each.

Pipeline Orchestration

Orchestration answers the question: who decides when pipelines run, in what order, and what happens when they fail?

Apache Airflow is the dominant open-source orchestrator. Pipelines are defined as DAGs (Directed Acyclic Graphs) in Python. Strong ecosystem, mature monitoring, complex to operate.

Prefect and Dagster are modern alternatives with better developer experience and native support for data-specific patterns (asset materialization, data lineage).

dbt Cloud handles orchestration for transformation workloads specifically.

At scale, orchestration complexity becomes a significant operational burden. Airbnb runs thousands of Airflow DAGs daily, managing DAG sprawl, dependency management, and failure recovery is a dedicated engineering function. CDOs should understand this cost when planning data engineering headcount.

Data pipeline reliability

Pipeline reliability is a first-order concern. A pipeline that produces wrong results silently is worse than a pipeline that fails loudly. Key practices:

  • Idempotency: Running a pipeline twice should produce the same result. Avoid append-only writes without deduplication.
  • Data quality checks at every stage: Don't wait until the end to validate, check at ingestion, at transformation, at load.
  • Alerting on quality, not just failure: A pipeline that runs but produces 40% null values has failed. Monitor completeness, freshness, and schema drift, not just job status.
  • Lineage documentation: Know which downstream consumers depend on each pipeline. When a pipeline changes, proactively notify downstream teams.

Practical architecture: the medallion pattern

The medallion architecture (Bronze → Silver → Gold) has become a de facto standard for lake and lakehouse implementations:

  • Bronze: Raw ingested data, unmodified, with metadata (load timestamp, source system). Never delete. This is your audit trail.
  • Silver: Cleaned, deduplicated, standardized data. Joins applied. Business rules partially enforced. Suitable for data science exploration.
  • Gold: Business-ready aggregated data. Optimized for specific analytical use cases. Used by BI tools and business stakeholders.

Uber, Netflix, and Databricks customers widely use this pattern. It provides clear data quality expectations at each layer and makes debugging faster, when a business metric is wrong, you know which layer to investigate first.

Quiz Questions

  1. Pourquoi l'ELT est-il devenu le pattern moderne par défaut ?

A) Il est plus rapide que l'ETL

B) Les warehouses cloud offrent du compute scalable, permettant de conserver la donnée brute et de re-transformer à la demande

C) Il nécessite moins de compétences techniques

D) Il réduit les coûts de stockage

Réponse: B

  1. Dans l'architecture Medallion, que contient la couche Bronze ?

A) Des données agrégées optimisées pour le BI

B) Des données nettoyées et standardisées

C) Des données brutes non modifiées avec métadonnées de chargement

D) Des données de production en temps réel

Réponse: C

  1. Quelle est la différence clé entre le streaming et le micro-batch ?

A) Le micro-batch est toujours plus lent que le batch

B) Le streaming traite chaque événement individuellement en temps réel, le micro-batch traite de petits lots toutes les quelques minutes

C) Le micro-batch est identique au streaming

D) Le streaming ne fonctionne qu'avec Kafka

Réponse: B

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Standardize pipelines on the Medallion bronze-silver-gold layering pattern
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.