Data contracts: the new standard for quality agreements between teams
Your data pipelines are breaking. Not because of bugs. Because nobody agreed on what the data should look like before it was built.
A data contractdata contractA formal agreement between the team that produces data and the teams that use it, defining structure, meaning, quality and who is accountable when it breaks.View full definition → is the solution, and it's one of the most practical governance innovations of the last five years.
The problem data contracts solve
In a typical data organization, the journey of data from source to consumer looks like this:
- A production engineer adds a new field to the checkout APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition → response
- They forget (or don't know) that a downstream data pipelinedata pipelineETL (Extract, Transform, Load) is a data integration process that pulls data from sources, reshapes it into a consistent format, and writes it into a target system.View full definition → depends on the existing schemaschemaA schema is the formal blueprint that defines how data is structured, named, typed, and related within a database, file, or message.View full definition →
- The pipelinepipelineAll active sales opportunities across the stages of the sales process, together with their combined potential value and probability of closing.View full definition → breaks silently, it processes the data but produces incorrect results
- Three weeks later, a business analyst notices revenue numbers look wrong
- The data team spends two days debugging across six systems to find the source
This happens everywhere, every week, at every organization with a non-trivial data infrastructure. It's not a people problem. It's an architecture problem: there is no formal agreement between data producers (engineering teams that build source systems) and data consumers (data teams that build pipelines and analytics).
A data contract is that agreement. It specifies:
- Schema: The structure of the data, field names, types, nullable status
- Semantics: What each field means, business definitions, not just technical names
- Quality SLAs: Minimum quality thresholds, completeness, freshness, validity
- Ownership: Who is responsible for the data (producer) and who is consuming it
- Change management: How and when schema changes will be communicated
When the production engineer adds that new field, the data contract is updated. The change triggers automated tests. The pipeline is updated before it breaks. The analyst's dashboard never sees the problem.
Data Contracts: The Missing Piece of the Data Puzzle
Knowledge check
1. According to the lesson, what is the fundamental root cause of the pipeline-breaking scenario described (engineer adds a field, pipeline breaks silently)?
2. In the context of data contracts, what does the 'Semantics' element specify that 'Schema' does not?
3. Why does the lesson recommend storing a data contract as a YAML or JSON file in version control (Git) alongside the producing code?
4. Select ALL elements that a data contract typically specifies, according to the lesson.
Select all the correct answers.
5. Select ALL statements that correctly describe how a data contract prevents the silent-breakage scenario.
Select all the correct answers.
Anatomy of a data contract
A data contract is typically a YAML or JSON file stored in version control (Git), alongside the code that produces the data. A simplified example would specify:
- contract.name: checkout_events
- contract.owner: ecommerce-platform-team
- contract.consumers: data-analytics-team, marketing-team
- contract.schema: order_id (string, required), order_amount_eur (decimal, required, "Total in EUR incl. VAT excl. shipping")
- contract.quality_sla: completeness 99.5%, freshness ≤15 min, no null order_id values
- contract.change_policy: 14-day notice, breaking changes require consumer_approval
When this contract is violated, order_amount_eur has null values, or freshness exceeds 15 minutes, an alert fires. The producer team is notified. Consumers aren't surprised by degraded data quality; they're notified of the violation.
Shopify's internal data contracts program
Shopify implemented an internal data contracts program at scale across their data platform. Their approach: any dataset served through their internal data platform must have a contract. Teams that consume data without a contract cannot hold producers accountable for quality. Teams that produce data without a contract cannot claim downstream consumers are using it correctly.
The outcome: a significant reduction in pipeline incidents caused by schema drift, a clearer accountability model for data qualitydata qualityThe degree to which data is fit for purpose: accurate, complete, consistent, timely, valid and unique. Poor quality data undermines analytics, reporting and AI.View full definition →, and faster root-cause analysis when issues do occur.
Implementing data contracts: where to start
Don't try to contract everything at once. Identify the five most business-critical data flows, the ones where a failure causes the most pain. Contract those first. Use the experience to refine your contract template and your change management process. Then scale.
Tools in this space: Soda (contract testing), Great Expectations (data quality validation), Gable.ai (data contract management platform), dbt contracts (built-in to dbt for transformation layer). The tooling is maturing rapidly, choose based on where in your stack you want to enforce contracts.
What to do, from this lesson
These actions are compiled in the role's Playbook.
- Implement data contracts on the five most business-critical data flows first
Related articles
Recent articles from the blog that build on this lesson.
- DataAirbus lost a quarter to misaligned revenue metrics: here is the playbook that prevents itWhen finance, sales, and product each calculate "revenue" differently, the damage shows up in board decks, budget fights, and delayed decisions. This playbook walks CDOs through building a semantic layer that makes metric definitions a shared organizational fact, not a tribal negotiation.
- DataThe modern ELT stack: how dbt, ingestion, and orchestration actually fit togetherThe ELT pattern has reshaped how data teams build pipelines, but the acronym hides considerable complexity in practice. This article breaks down how dbt, ingestion tools, and orchestration layers interact, and where the real architectural decisions lie.
- DataHow JPMorgan Chase built data contracts across 50+ domainsJPMorgan Chase spent years grappling with fragmented data ownership across hundreds of business lines before systematically formalizing who owns what and on what terms. Their approach to data contracts offers a working model for CDOs who need accountability without organizational paralysis.