Modern data architecture
Design data warehouses, lakes, lakehouses, data mesh, pipelines and a modern data stack
Every data strategy dies or thrives on its architecture. You can hire the best data scientists in the market, but if your storage, pipelines and transformation layers are a tangle of legacy compromises, your teams spend their days firefighting instead of delivering. This block puts the architecture decisions where they belong, on your desk.
You start with the foundations that everyone argues about and few get right. Data Warehouse, Lake or Lakehouse, and how to choose without following the vendor of the month. Cloud infrastructure with real cost math, not marketing brochures. Pipelines that actually hold up, from ETL and ELT to batch, streaming and the Medallion architecture that keeps your data trustworthy at every stage.
Then you go where the sharp CDOs are looking. Data Mesh, its real principles, the conditions it demands and the honest critiques that consultants skip. Data Products treated as products, with design, ownership and lifecycle management. Streaming with Apache Kafka for the moments when yesterday's data is worthless.
Finally you industrialize. dbt turns SQL transformation from craft into factory. Data Observability catches problems before your users open a ticket. Performance and scalability work, partitioning, clustering and cost optimization, makes sure your platform grows without your invoice growing faster.
This is not a course for engineers. It is the architecture fluency that lets you challenge your teams, kill bad proposals early and fund the right ones. You will not write the pipelines. You will know exactly why one design wins and another bleeds cash and credibility. That is the difference between a CDO who approves architecture and one who owns it.
Ce que vous allez maîtriser
- Choose between Warehouse, Lake and Lakehouse based on your actual workloads and costs
- Evaluate cloud data infrastructure with a clear grip on services, migration risk and total cost
- Design pipelines across batch and streaming using the Medallion architecture
- Assess whether Data Mesh fits your organization or sets it up to fail
- Treat data as products with real ownership, design and lifecycle governance
- Deploy dbt and observability so quality problems surface before your users do
- Optimize performance and cost through partitioning, clustering and smart scaling
Modules
Compare storage architectures, cloud services and data pipeline types
Master data mesh principles, data product design and real-time streaming
Industrialize transformations with dbt, monitor data quality, and optimize performance and cost
Compares batch and streaming paradigms, event-driven platforms, and lakehouse designs that unify analytics and machine learning on shared data.
Shows how to catalog data assets, trace lineage for impact analysis, and control cloud data costs through FinOps practices.