Pipelines de données : ETL/ELT, batch, streaming et architecture Medallion
Data pipelines are the circulatory system of your data platform. They move data from where it's produced to where it's consumed. When they work, nobody notices. When they fail, everything stops.
A CDO who doesn't understand data pipelinedata pipelineETL (Extract, Transform, Load) is a data integration process that pulls data from sources, reshapes it into a consistent format, and writes it into a target system.View full definition → architecture can't diagnose failures, can't have credible conversations with engineering teams, and can't make good architectural decisions.
ETL vs. ELTELTELT (Extract, Load, Transform) is a data integration pattern where raw data is loaded into a target system first, then transformed inside it using the platform's compute power.View full definition →: the fundamental pattern shift
Traditional ETL (Extract, Transform, Load) transforms data before loading it to the destination. ELT (Extract, Load, Transform) loads raw data first, then transforms it inside the destination system.
ETL made sense when compute was expensive and storage was cheap, you minimized what you stored. ELT makes sense when cloud warehouses (Snowflake, BigQuery) offer scalable compute, and you want to preserve raw data for flexible reuse.
Modern default: ELT. Load raw data to your warehouse or lake. Transform using SQLSQLSales Qualified Lead: a prospect the sales team has validated as ready for direct outreach and a proposal, having passed clear qualification criteria.View full definition → (dbt) or Spark. Re-transform as needs evolve without reloading source data.
ETL vs ELT: Key Differences Explained
Knowledge check
1. What is the fundamental distinction between ETL and ELT?
2. Why has ELT become the modern default over traditional ETL?
3. A fraud detection system must react to transactions as they happen. Which processing pattern is most appropriate?
4. Select ALL correct statements about batch, streaming, and micro-batch processing.
Select all the correct answers.
5. Select ALL correct statements about pipeline orchestration and Apache Airflow.
Select all the correct answers.
Batch vs. streaming: the latency spectrum
Batch processing moves data in scheduled windows, hourly, daily, weekly. Tools: Apache Spark, dbt, AWS Glue. Use when: latency is acceptable (overnight reports, weekly models), data volume is high, and processing is complex.
Streaming processing moves data in real-time or near-real-time (seconds to minutes). Tools: Apache Kafka, Apache Flink, AWS Kinesis, Google Pub/Sub. Use when: business decisions depend on current data (fraud detection, real-time personalization, operational dashboards).
Micro-batch is the middle ground, small batches processed every few minutes. Apache Spark Structured Streaming supports this. Often sufficient when true real-time isn't required but hourly batch is too slow.
Most organizations need both: batch for heavy analytical workloads, streaming for operational use cases. Architect to support both patterns, don't build separate, incompatible systems for each.
PipelinePipelineAll active sales opportunities across the stages of the sales process, together with their combined potential value and probability of closing.View full definition → Orchestration
Orchestration answers the question: who decides when pipelines run, in what order, and what happens when they fail?
Apache Airflow is the dominant open-source orchestrator. Pipelines are defined as DAGs (Directed Acyclic Graphs) in Python. Strong ecosystem, mature monitoring, complex to operate.
Prefect and Dagster are modern alternatives with better developer experience and native support for data-specific patterns (asset materialization, data lineagedata lineageData lineage maps how data moves and transforms across systems, from origin to consumption, showing where it came from, what changed it, and where it goes.View full definition →).
dbt Cloud handles orchestration for transformation workloads specifically.
At scale, orchestration complexity becomes a significant operational burden. Airbnb runs thousands of Airflow DAGs daily, managing DAG sprawl, dependency management, and failure recovery is a dedicated engineering function. CDOs should understand this cost when planning data engineering headcount.
Data pipeline reliability
Pipeline reliability is a first-order concern. A pipeline that produces wrong results silently is worse than a pipeline that fails loudly. Key practices:
- Idempotency: Running a pipeline twice should produce the same result. Avoid append-only writes without deduplication.
- Data quality checks at every stage: Don't wait until the end to validate, check at ingestion, at transformation, at load.
- Alerting on quality, not just failure: A pipeline that runs but produces 40% null values has failed. Monitor completeness, freshness, and schemaschemaA schema is the formal blueprint that defines how data is structured, named, typed, and related within a database, file, or message.View full definition → drift, not just job status.
- Lineage documentation: Know which downstream consumers depend on each pipeline. When a pipeline changes, proactively notify downstream teams.
Practical architecture: the medallion pattern
The medallion architecture (Bronze → Silver → Gold) has become a de facto standard for lake and lakehouse implementations:
- Bronze: Raw ingested data, unmodified, with metadata (load timestamp, source system). Never delete. This is your audit trail.
- Silver: Cleaned, deduplicated, standardized data. Joins applied. Business rules partially enforced. Suitable for data science exploration.
- Gold: Business-ready aggregated data. Optimized for specific analytical use cases. Used by BIBITechnologies and processes that turn raw data into actionable insights via reporting, dashboards and analysis, so teams can decide based on facts rather than intuition.View full definition → tools and business stakeholders.
Uber, Netflix, and Databricks customers widely use this pattern. It provides clear data qualitydata qualityThe degree to which data is fit for purpose: accurate, complete, consistent, timely, valid and unique. Poor quality data undermines analytics, reporting and AI.View full definition → expectations at each layer and makes debugging faster, when a business metric is wrong, you know which layer to investigate first.
Quiz Questions
- Pourquoi l'ELT est-il devenu le pattern moderne par défaut ?
A) Il est plus rapide que l'ETL
B) Les warehouses cloud offrent du compute scalable, permettant de conserver la donnée brute et de re-transformer à la demande
C) Il nécessite moins de compétences techniques
D) Il réduit les coûts de stockage
Réponse: B
- Dans l'architecture Medallion, que contient la couche Bronze ?
A) Des données agrégées optimisées pour le BI
B) Des données nettoyées et standardisées
C) Des données brutes non modifiées avec métadonnées de chargement
D) Des données de production en temps réel
Réponse: C
- Quelle est la différence clé entre le streaming et le micro-batch ?
A) Le micro-batch est toujours plus lent que le batch
B) Le streaming traite chaque événement individuellement en temps réel, le micro-batch traite de petits lots toutes les quelques minutes
C) Le micro-batch est identique au streaming
D) Le streaming ne fonctionne qu'avec Kafka
Réponse: B
What to do, from this lesson
These actions are compiled in the role's Playbook.
- Standardize pipelines on the Medallion bronze-silver-gold layering pattern
Related articles
Recent articles from the blog that build on this lesson.
- DataThe modern ELT stack: how dbt, ingestion, and orchestration actually fit togetherThe ELT pattern has reshaped how data teams build pipelines, but the acronym hides considerable complexity in practice. This article breaks down how dbt, ingestion tools, and orchestration layers interact, and where the real architectural decisions lie.
- DataHow Cloudflare rebuilt its data stack around dbt, Fivetran, and AirflowCloudflare's rapid growth exposed the limits of hand-coded SQL pipelines and fragmented ingestion scripts that no engineer wanted to touch. This case study traces how the company restructured its analytical data layer using a modern ELT approach, and what that shift actually required in practice.
- DataThe modern ELT stack explained: dbt, ingestion, and orchestration working togetherThe shift from ETL to ELT reshaped how data teams build pipelines, but the real complexity lies in understanding how the three layers, ingestion, transformation, and orchestration, actually interact. This article breaks down the mechanics of the modern stack with concrete examples, and explains where the genuine tradeoffs sit for leaders making architecture decisions.