+45 XP

Shift-left data quality: embedding governance in the engineering pipeline

Data quality failures have a geography: they almost always originate at the source.

A report shows wrong revenue numbers. The investigation traces back through the Data warehouse, the ETL pipeline, the staging database, the API integration, and finally to a source system that started sending malformed data six weeks ago. Six weeks of bad data in production. Six weeks of decisions made on incorrect information.

The "shift-left" principle borrows from software engineering: catch defects as early as possible in the development process, because fixing a bug in production is 100x more expensive than catching it in code review. Applied to data: catch quality issues at the source, not after the fact.

What shift-left data quality looks like

At the source system: Validation rules built into the application that produces the data. If a field cannot be null, the application enforces it, the Data pipeline never sees null values because they never enter the system.

At the ingestion layer: Data quality checks run immediately when data enters your infrastructure. If the schema doesn't match the contract, the pipeline stops. If completeness drops below the SLA, an alert fires. Great Expectations and Soda run assertion-style checks here (Great Expectations moved to its GX 1.0 API in 2024, so older Expectation syntax from tutorials may not match current docs). Monte Carlo works alongside these tools but plays a different role: it is data observability, monitoring freshness, volume and schema drift across your warehouse and alerting when something looks off, rather than blocking a pipeline on a fixed rule. Note that Monte Carlo acquired the commercial side of Great Expectations in 2024, though the open-source project continues.

At the transformation layer: dbt (data build tool) has built-in testing: not-null tests, unique tests, referential integrity tests, accepted-value tests, custom SQL tests. Every dbt model should have tests. A dbt run that includes failing tests should not deploy to production. Many organizations run their dbt tests in CI/CD pipelines, no untested transformation reaches the Data warehouse.

At the serving layer: Dashboards and reports that expose data to business users should include data quality indicators: "Last refreshed: 2 hours ago. Quality score: 94%. Known issues: 0."

Data Contracts: The Key to Data Quality - with Chad Sanderson

Watch on YouTube

Knowledge check

1. What is the core principle behind 'shift-left' data quality?

2. Why does the lesson argue that catching a data quality issue at the source is preferable to catching it in a production report?

3. In the transformation layer, what is the recommended best practice regarding dbt tests and deployment?

MULTIPLE CHOICE

4. Select ALL statements that correctly describe where shift-left data quality controls operate.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL practices that reflect building data quality into a CI/CD pipeline.

Select all the correct answers.

Building data quality into CI/CD

The gold standard for shift-left data quality: data quality tests run in CI/CD pipelines, failing builds are blocked from deployment, and quality metrics are tracked in the same dashboard as engineering metrics.

This requires:

  1. Test coverage for data: Every critical pipeline has documented quality tests. Tracked as a metric: "Percentage of data assets with quality tests: 67%." The CDO should set a target, say, 90%, and track it quarterly.
  1. Automated validation on merge: When a data engineer submits a pull request that changes a pipeline, automated tests run against a sample of production data. Schema changes that would break downstream contracts fail the build before merge.
  1. Quality gates for promotion: Data doesn't move from staging to production without passing quality checks. This is standard in software engineering (you don't deploy broken code). It should be standard in data engineering too.

The Airbnb Minerva framework

Airbnb built an internal framework called Minerva to solve a specific problem: hundreds of analysts were defining the same metrics differently, creating inconsistency that undermined trust in data.

Minerva is a metrics layer, a central repository where business metrics are defined once (by the business, with the data team), and consumed consistently across all BI tools, data science models, and experiments.

The key insight: Minerva shifts "what does this metric mean?" from ad hoc analyst judgment to a governed, version-controlled definition. When a business metric changes (a new return policy changes how revenue is counted), the definition is updated in one place and flows to all consumers.

This is shift-left data quality applied to semantic consistency rather than technical quality. It's one of the highest-leverage data governance investments Airbnb has made, and the pattern is now widespread. The dbt Semantic Layer runs on MetricFlow (dbt Labs acquired Transform and its MetricFlow engine in 2023 and folded it in), Looker has its measures, and Cube offers a standalone metrics layer. All solve the same problem Minerva did.

What CDOs should track

  • Test coverage: % of data assets with automated quality tests
  • Data pipeline incident rate: Number of pipeline failures per week, trending
  • Mean time to detection (MTTD): How quickly data quality issues are detected
  • Mean time to resolution (MTTR): How quickly detected issues are resolved
  • Data freshness SLA compliance: % of datasets meeting their freshness SLA

Tracked monthly, these metrics tell you whether your shift-left program is working.

Key Takeaways

  • Bad data usually enters at the source, so the cheapest place to catch it is upstream, not in the warehouse.
  • Use the right tool for each layer: assertion checks (Great Expectations on its GX 1.0 API, Soda) catch known rule violations, while observability tools (Monte Carlo) surface anomalies you did not think to test for.
  • dbt tests belong in CI/CD, with failing builds blocked from production, the same discipline software teams already apply to code.
  • A metrics layer (Airbnb's Minerva, the dbt Semantic Layer on MetricFlow, Looker, Cube) governs what a metric means, so revenue is counted the same way everywhere.
  • Track MTTD, MTTR, test coverage, incident rate and freshness SLA compliance monthly to know if the program is actually working.

What to do, from this lesson

These actions are compiled in the role's Playbook.

  • Implement data contracts on the five most business-critical data flows first
See the full action playbook →

Related articles

Recent articles from the blog that build on this lesson.