Block 3

Modern data architecture

Design data warehouses, lakes, lakehouses, data mesh, pipelines and a modern data stack

5 Modules·15 Lessons

Every data strategy dies or thrives on its architecture. You can hire the best data scientists in the market, but if your storage, pipelines and transformation layers are a tangle of legacy compromises, your teams spend their days firefighting instead of delivering. This block puts the architecture decisions where they belong, on your desk.

You start with the foundations that everyone argues about and few get right. Data Warehouse, Lake or Lakehouse, and how to choose without following the vendor of the month. Cloud infrastructure with real cost math, not marketing brochures. Pipelines that actually hold up, from ETL and ELT to batch, streaming and the Medallion architecture that keeps your data trustworthy at every stage.

Then you go where the sharp CDOs are looking. Data Mesh, its real principles, the conditions it demands and the honest critiques that consultants skip. Data Products treated as products, with design, ownership and lifecycle management. Streaming with Apache Kafka for the moments when yesterday's data is worthless.

Finally you industrialize. dbt turns SQL transformation from craft into factory. Data Observability catches problems before your users open a ticket. Performance and scalability work, partitioning, clustering and cost optimization, makes sure your platform grows without your invoice growing faster.

This is not a course for engineers. It is the architecture fluency that lets you challenge your teams, kill bad proposals early and fund the right ones. You will not write the pipelines. You will know exactly why one design wins and another bleeds cash and credibility. That is the difference between a CDO who approves architecture and one who owns it.

What you'll master

  • Choose between Warehouse, Lake and Lakehouse based on your actual workloads and costs
  • Evaluate cloud data infrastructure with a clear grip on services, migration risk and total cost
  • Design pipelines across batch and streaming using the Medallion architecture
  • Assess whether Data Mesh fits your organization or sets it up to fail
  • Treat data as products with real ownership, design and lifecycle governance
  • Deploy dbt and observability so quality problems surface before your users do
  • Optimize performance and cost through partitioning, clustering and smart scaling

Modules

Frequently asked questions

Who is the Modern data architecture block for?

It targets CDOs, data leaders and executives who approve architecture decisions without writing the code. The 5 modules and 15 lessons cover warehouses, lakes, lakehouses, pipelines, Data Mesh, dbt and observability at the level of decision-making, not implementation. If you need to challenge a vendor proposal or a team's design, this is the right level.

Do I need to be technical to follow it?

No. Modern data architecture assumes you will never build the pipelines yourself. The goal is architecture fluency: understanding why one design wins and another bleeds cash, so you can kill weak proposals early and fund the right ones. Some familiarity with SQL and cloud vocabulary helps but is not required.

What is the difference between a data warehouse, a data lake and a lakehouse?

A warehouse stores structured, modeled data for analytics; a lake stores raw data of any format at low cost; a lakehouse combines lake storage with warehouse-style query and governance layers. The first lesson of the block walks through the choice based on your actual workloads and costs rather than the vendor of the month, and a later lesson covers how the lakehouse unifies analytics and machine learning on shared data.

Where should I start if my current data platform is a tangle of legacy pipelines?

Start with the storage and cloud architecture module: warehouse versus lake versus lakehouse, cloud services with real cost math and migration risk, then pipelines from ETL and ELT to batch, streaming and the Medallion architecture. Fixing the storage and transformation layers first is what stops teams from firefighting. Data Mesh and streaming come later, once the foundations hold.

Does the block take a position on Data Mesh?

Yes, a nuanced one. The lesson on Data Mesh covers its principles, the organizational conditions it demands and the honest criticisms consultants tend to skip, so you can judge whether it fits your organization or sets it up to fail. A companion lesson treats data products with real design, ownership and lifecycle management, which is where most Data Mesh attempts break down.

How does the block handle cloud data costs?

Two ways. The performance and scalability lesson covers partitioning, clustering and cost optimization so a growing platform does not grow the invoice faster, and a dedicated Data FinOps lesson addresses controlling cloud data spend. The cloud infrastructure lesson also frames services and migration in total cost terms rather than list prices.