Block 3

Modern data architecture

Design data warehouses, lakes, lakehouses, data mesh, pipelines and a modern data stack

5 Modules·15 Lessons

Every data strategy dies or thrives on its architecture. You can hire the best data scientists in the market, but if your storage, pipelines and transformation layers are a tangle of legacy compromises, your teams spend their days firefighting instead of delivering. This block puts the architecture decisions where they belong, on your desk.

You start with the foundations that everyone argues about and few get right. Data Warehouse, Lake or Lakehouse, and how to choose without following the vendor of the month. Cloud infrastructure with real cost math, not marketing brochures. Pipelines that actually hold up, from ETL and ELT to batch, streaming and the Medallion architecture that keeps your data trustworthy at every stage.

Then you go where the sharp CDOs are looking. Data Mesh, its real principles, the conditions it demands and the honest critiques that consultants skip. Data Products treated as products, with design, ownership and lifecycle management. Streaming with Apache Kafka for the moments when yesterday's data is worthless.

Finally you industrialize. dbt turns SQL transformation from craft into factory. Data Observability catches problems before your users open a ticket. Performance and scalability work, partitioning, clustering and cost optimization, makes sure your platform grows without your invoice growing faster.

This is not a course for engineers. It is the architecture fluency that lets you challenge your teams, kill bad proposals early and fund the right ones. You will not write the pipelines. You will know exactly why one design wins and another bleeds cash and credibility. That is the difference between a CDO who approves architecture and one who owns it.

What you'll master

  • Choose between Warehouse, Lake and Lakehouse based on your actual workloads and costs
  • Evaluate cloud data infrastructure with a clear grip on services, migration risk and total cost
  • Design pipelines across batch and streaming using the Medallion architecture
  • Assess whether Data Mesh fits your organization or sets it up to fail
  • Treat data as products with real ownership, design and lifecycle governance
  • Deploy dbt and observability so quality problems surface before your users do
  • Optimize performance and cost through partitioning, clustering and smart scaling

Modules