Shift-left data quality: embedding governance in the engineering pipeline
Data qualityData qualityThe degree to which data is fit for purpose: accurate, complete, consistent, timely, valid and unique. Poor quality data undermines analytics, reporting and AI.View full definition → failures have a geography: they almost always originate at the source.
A report shows wrong revenue numbers. The investigation traces back through the Data warehouseData warehouseA central repository that consolidates data from many source systems into a structured, query-optimized store designed for analytics, reporting, and business intelligence.View full definition →, the ETL pipelineETL pipelineAn automated sequence of steps that moves data from source to destination: ingestion, transformation, validation, and loading, so it arrives clean and ready to use.View full definition →, the staging database, the APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition → integration, and finally to a source system that started sending malformed data six weeks ago. Six weeks of bad data in production. Six weeks of decisions made on incorrect information.
The "shift-left" principle borrows from software engineering: catch defects as early as possible in the development process, because fixing a bug in production is 100x more expensive than catching it in code review. Applied to data: catch quality issues at the source, not after the fact.
What shift-left data quality looks like
At the source system: Validation rules built into the application that produces the data. If a field cannot be null, the application enforces it, the Data pipeline never sees null values because they never enter the system.
At the ingestion layer: Data quality checks run immediately when data enters your infrastructure. If the schema doesn't match the contract, the pipeline stops. If completeness drops below the SLA, an alert fires. Great Expectations and Soda run assertion-style checks here (Great Expectations moved to its GX 1.0 API in 2024, so older Expectation syntax from tutorials may not match current docs). Monte Carlo works alongside these tools but plays a different role: it is data observabilitydata observabilityThe practice of continuously monitoring the health of your data so you know when it breaks, drifts or goes missing before it reaches a report or a model.View full definition →, monitoring freshness, volume and schemaschemaA schema is the formal blueprint that defines how data is structured, named, typed, and related within a database, file, or message.View full definition → drift across your warehouse and alerting when something looks off, rather than blocking a pipelinepipelineAll active sales opportunities across the stages of the sales process, together with their combined potential value and probability of closing.View full definition → on a fixed rule. Note that Monte Carlo acquired the commercial side of Great Expectations in 2024, though the open-source project continues.
At the transformation layer: dbt (data build tool) has built-in testing: not-null tests, unique tests, referential integrity tests, accepted-value tests, custom SQLSQLSales Qualified Lead: a prospect the sales team has validated as ready for direct outreach and a proposal, having passed clear qualification criteria.View full definition → tests. Every dbt model should have tests. A dbt run that includes failing tests should not deploy to production. Many organizations run their dbt tests in CI/CD pipelines, no untested transformation reaches the Data warehouse.
At the serving layer: Dashboards and reports that expose data to business users should include data quality indicators: "Last refreshed: 2 hours ago. Quality score: 94%. Known issues: 0."
Data Contracts: The Key to Data Quality - with Chad Sanderson
Knowledge check
1. What is the core principle behind 'shift-left' data quality?
2. Why does the lesson argue that catching a data quality issue at the source is preferable to catching it in a production report?
3. In the transformation layer, what is the recommended best practice regarding dbt tests and deployment?
4. Select ALL statements that correctly describe where shift-left data quality controls operate.
Select all the correct answers.
5. Select ALL practices that reflect building data quality into a CI/CD pipeline.
Select all the correct answers.
Building data quality into CI/CD
The gold standard for shift-left data quality: data quality tests run in CI/CD pipelines, failing builds are blocked from deployment, and quality metrics are tracked in the same dashboard as engineering metrics.
This requires:
- Test coverage for data: Every critical pipeline has documented quality tests. Tracked as a metric: "Percentage of data assets with quality tests: 67%." The CDO should set a target, say, 90%, and track it quarterly.
- Automated validation on merge: When a data engineer submits a pull request that changes a pipeline, automated tests run against a sample of production data. Schema changes that would break downstream contracts fail the build before merge.
- Quality gates for promotion: Data doesn't move from staging to production without passing quality checks. This is standard in software engineering (you don't deploy broken code). It should be standard in data engineering too.
The Airbnb Minerva framework
Airbnb built an internal framework called Minerva to solve a specific problem: hundreds of analysts were defining the same metrics differently, creating inconsistency that undermined trust in data.
Minerva is a metrics layer, a central repository where business metrics are defined once (by the business, with the data team), and consumed consistently across all BI tools, data science models, and experiments.
The key insight: Minerva shifts "what does this metric mean?" from ad hoc analyst judgment to a governed, version-controlled definition. When a business metric changes (a new return policy changes how revenue is counted), the definition is updated in one place and flows to all consumers.
This is shift-left data quality applied to semantic consistency rather than technical quality. It's one of the highest-leverage data governance investments Airbnb has made, and the pattern is now widespread. The dbt Semantic Layer runs on MetricFlow (dbt Labs acquired Transform and its MetricFlow engine in 2023 and folded it in), Looker has its measures, and Cube offers a standalone metrics layer. All solve the same problem Minerva did.
What CDOs should track
- Test coverage: % of data assets with automated quality tests
- Data pipeline incident rate: Number of pipeline failures per week, trending
- Mean time to detection (MTTD): How quickly data quality issues are detected
- Mean time to resolution (MTTR): How quickly detected issues are resolved
- Data freshness SLA compliance: % of datasets meeting their freshness SLA
Tracked monthly, these metrics tell you whether your shift-left program is working.
Key Takeaways
- Bad data usually enters at the source, so the cheapest place to catch it is upstream, not in the warehouse.
- Use the right tool for each layer: assertion checks (Great Expectations on its GX 1.0 API, Soda) catch known rule violations, while observability tools (Monte Carlo) surface anomalies you did not think to test for.
- dbt tests belong in CI/CD, with failing builds blocked from production, the same discipline software teams already apply to code.
- A metrics layer (Airbnb's Minerva, the dbt Semantic Layer on MetricFlow, Looker, Cube) governs what a metric means, so revenue is counted the same way everywhere.
- Track MTTD, MTTR, test coverage, incident rate and freshness SLA compliance monthly to know if the program is actually working.
What to do, from this lesson
These actions are compiled in the role's Playbook.
- Implement data contracts on the five most business-critical data flows first
Related articles
Recent articles from the blog that build on this lesson.
- DataAI slop detectors make your training data worse before they make it betterFiltering AI-generated text from training datasets sounds like straightforward hygiene, but a recent experiment shows the cure can degrade model performance more than the contamination itself. CDOs who treat this as a tooling problem will miss the governance question underneath it.
- DataData observability: catching bad data before it reaches decisionsBad data doesn't announce itself. This playbook shows CDOs how to build detection mechanisms that intercept data quality failures before they corrupt reports, models, and the decisions that follow.
- DataHow Cloudflare rebuilt its data stack around dbt, Fivetran, and AirflowCloudflare's rapid growth exposed the limits of hand-coded SQL pipelines and fragmented ingestion scripts that no engineer wanted to touch. This case study traces how the company restructured its analytical data layer using a modern ELT approach, and what that shift actually required in practice.