DataData Governance

How JPMorgan Chase built data contracts across 50+ domains

JPMorgan Chase spent years wrestling with data inconsistencies across hundreds of business lines before committing to a structured data contract framework. The mechanics they chose, and the organizational friction they encountered, offer a practical blueprint for CDOs facing the same ownership vacuum.

🎙️

Listen to the podcast

4 min

By the early 2020s, JPMorgan Chase was operating with over 50 distinct data domains, each managed by separate lines of business: retail banking, investment banking, asset management, treasury services, and more. The firm's data engineering teams routinely discovered that the same field, say "customer risk rating," carried different definitions, refresh cadences, and lineage documentation depending on which domain produced it. Downstream consumers, whether risk analysts, compliance officers, or quantitative trading desks, had no reliable way to know whether the data they were reading had been validated, who owned it, or what transformations it had passed through. This is not a JPMorgan-specific pathology; it is the default state of any large financial institution that has grown through acquisitions and decentralized technology decisions. What distinguishes JPMorgan's trajectory is the deliberate, enterprise-scale response they committed to.

What they did

The firm's data strategy, publicly discussed by former CDO Teresa Heitsenrether and her successors in various industry forums, centered on treating data as a product with explicit contractual obligations between producers and consumers. This means a data contract in their context is not a legal document; it is a formal specification that defines schema, semantics, quality thresholds, SLA-level availability commitments, and the named domain owner accountable for all of it.

JPMorgan built this out through several concrete mechanisms.

They established a federated data governance model with a central Data Management Office setting standards, while domain stewards within each line of business held operational accountability. The CDO function did not try to centralize data production; it standardized the contract layer on top of distributed ownership. Each domain team was required to publish a data contract before any dataset could be registered in the firm's internal data catalog, which by the mid-2020s was running on a combination of internally developed tooling and commercial catalog vendors (the firm has publicly referenced partnerships with vendors including Collibra; as a commercial catalog vendor, Collibra's own claims about adoption and outcomes should be read with that context in mind).

The contract specification itself required four things: a canonical schema with documented field-level definitions, data quality rules expressed as executable tests, an identified data product owner with a named backup, and a published SLA covering freshness and availability. Teams that failed to maintain SLA compliance faced escalation to domain leadership, not to a central data team. That organizational detail matters. By placing accountability at the domain level rather than in a shared data engineering function, JPMorgan avoided the common failure mode where central teams become the bottleneck and the scapegoat simultaneously.

They also introduced what internal documentation and conference presentations described as "data lineage attestation," where domain owners periodically certify that upstream dependencies have been reviewed and that any breaking changes to source systems will trigger a contract renegotiation process. This brought version control discipline to data pipelines that had previously changed without notice to downstream teams.

The role of incentives

One structural decision that made this work was tying data contract compliance to the performance review criteria for domain data stewards. This is mentioned in the context of JPMorgan's broader data literacy program, which the firm reported had trained tens of thousands of employees in data fluency by 2023, a figure cited in public communications from the firm itself and therefore worth treating as directional rather than independently audited. The point is that contract ownership became a job responsibility with consequences, not a best practice documented in a wiki nobody reads.

The results

Exact operational metrics from JPMorgan's internal data contract rollout are not publicly disclosed, and any specific figures circulating from vendor case studies should be treated accordingly. What is documented through regulatory filings, earnings calls, and industry conference presentations:

The firm has consistently cited reductions in data-related incident resolution time as one output of clearer ownership structures. When a data quality issue surfaces, the contract identifies the owner immediately, which compresses the triage phase from days to hours in documented internal reviews.

Regulatory compliance work has benefited in measurable terms. JPMorgan operates under intense scrutiny from the OCC, Federal Reserve, and international regulators who increasingly require firms to demonstrate data lineage for risk models. The ability to point auditors to a contract that documents provenance, transformation logic, and current ownership has reduced the manual effort involved in regulatory data requests. The firm has not published a precise hours-saved figure for this, and any vendor claiming to quantify it on JPMorgan's behalf should be read skeptically.

Cross-domain data reuse has increased. Teams that previously rebuilt similar datasets independently because they could not trust an existing source have shifted toward consuming certified data products. This reduces duplicated engineering work, though the firm has not published a cost-of-duplication baseline against which to measure the improvement.

What transfers

The JPMorgan case is instructive precisely because their scale is not a prerequisite for the approach. The core logic transfers to any organization with more than three or four data-producing domains.

The first transferable principle is separating governance standards from data production. A central CDO office that tries to own both will fail at one or both. The contract layer is the right place for central authority; production stays federated.

The second is making ownership named and individual, not assigned to a team or a system. Teams rotate, systems get renamed, but a named person with a named backup creates a durable accountability thread.

Third, contracts need teeth. At JPMorgan, compliance linked to performance management. In smaller organizations without that HR lever, the equivalent might be blocking dataset registration in the catalog until a valid contract exists, or requiring a signed contract before a dataset can be consumed by a regulated process.

Where context differs: JPMorgan's regulatory environment created external pressure that accelerated internal adoption. If your organization lacks that external forcing function, expect the change management work to take longer and require more visible executive sponsorship. Organizations in less regulated industries also have more flexibility in how they define contract scope, which can be an advantage or an excuse to under-specify.

The practical starting point for a CDO is not to mandate contracts across every domain simultaneously. JPMorgan's rollout was phased, beginning with the highest-risk and most-consumed datasets. Identify the five datasets that cause the most downstream pain when they break, assign contracts to those first, and use the early wins to build the organizational muscle before scaling.

Data contracts do not eliminate data quality problems. They create a clear accountability structure so that problems surface faster and ownership disputes become shorter. That is a meaningful operational improvement, and it is the right framing to use when building the business case internally.

Finished reading?

Validate your read to earn XP and feed your radar.