The shared data infrastructure problem that quietly breaks every cross-agency program
When two agencies cannot agree on what a "household" means, no amount of technology fixes the mismatch. This article explains how interoperability standards actually work in government data environments and where the traps are for CDOs who underestimate the governance layer.
Claude VectorData & Analytics LeadSeptember 29, 2026Interoperability is the word that appears in every federal IT strategy document and almost every state digital transformation plan. It also appears, without resolution, in the post-mortems of programs that failed to share data across agencies despite years of effort. The concept itself is not complicated. The execution is, and the gap between those two things costs governments real program outcomes.
The specific confusion worth naming: most CDOs entering the public sector think interoperability is primarily a technical problem. Pick a standard, build an APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition →, done. In practice, the technical layer is the part that gets resolved first. What breaks programs is the governance layer underneath: who owns the canonical definition of a field, who can authorize a data-sharing agreement, which agency bears liability when a record is wrong, and how you handle the fact that Agency A has been collecting "date of birth" as a free-text field since 1998.
Why cross-agency data failure lands on the CDO's desk, not the CIO's
In a private company, a data integration failure is a cost problem. In a government agency, it is a service delivery failure with names attached. When the U.S. Department of Veterans Affairs and the Social Security Administration cannot reconcile beneficiary records, veterans lose benefits they are entitled to. When child welfare and public health systems share no common client identifier, a child can fall through both at once. The CDO is accountable for that outcome in a way the CIO is not, because the CIO owns the pipes and the CDO owns what flows through them.
Federal procurement law compounds this. Under the Federal Acquisition Regulation, data-sharing arrangements between agencies often require formal interagency agreements, sometimes Economy Act orders, sometimes memoranda of understanding that take six to eighteen months to execute. A CDO who commits to a cross-agency data initiative in a budget cycle without accounting for that procurement timeline will miss it. Auditors from the Government Accountability Office check whether data-sharing agreements actually authorize the use described in a program's documentation. If they do not, the finding goes in the report and the CDO explains it to the appropriations subcommittee.
How public sector data interoperability actually works: the mechanics
The architecture has three layers, and they have to be addressed in sequence or the whole thing unravels.
The first layer is semantic: agreeing on what a term means before you agree on how to transmit it. The U.S. government has the National Information Exchange Model (NIEM), a data dictionary that defines common elements across justice, health, immigration, and emergency management. NIEM does not mandate technology. It mandates meaning. When a county sheriff's department and a state health department both describe an "incident" using NIEM-conformant definitions, they can exchange records without a custom translation table. The European Union's approach through the Joinup interoperability framework serves a similar function across member state administrations.
The second layer is technical: APIs, data formats, and transport protocols. Most U.S. federal health programs now reference HL7 FHIR as the required standard for health data exchange, a mandate formalized through the 21st Century Cures Act. State Medicaid agencies implementing FHIR-based APIs have found that the standard itself is not the bottleneck; the bottleneck is that legacy MMIS systems were not built to expose FHIR-conformant endpoints, and replacing them costs more than the API work by an order of magnitude.
The third layer is legal and governance: data-sharing agreements that specify who can access what, under which statutory authority, for which purpose, and with which audit trail. This is where most cross-agency programs stall. Each agency has its own privacy counsel, its own interpretation of the Privacy Act of 1974, and its own risk tolerance. The CDO's job is to drive those conversations to resolution, not to assume legal will handle it independently.
A concrete example: the California Health and Human Services Agency's Child Welfare Data LakeData LakeA data lake is a centralized repository that stores large volumes of raw data in its native format, from structured tables to unstructured files, until needed.View full definition →, stood up between 2021 and 2024, required NIEM-based semantic alignment across county probation, state foster care, and Medi-Cal records before a single API was written. The semantic work took fourteen months. The technical integration took four.
When shared infrastructure works and when it becomes an expensive liability
Shared data infrastructure between agencies makes sense when programs genuinely share clients across the full lifecycle, not just at a single point of contact. Workforce development, housing assistance, and public health are good candidates because the same individuals cycle through all three systems over years. Building a shared identifier and a shared data model pays back across the program lifetime.
It does not make sense when the data-sharing rationale is "it would be interesting to combine these." Understanding the full shape of what public data exists before committing to a shared infrastructure build often reveals that administrative records already answer the question without requiring a new integration layer.
The honest tradeoffs CDOs should name explicitly before committing to shared infrastructure:
- Governance overhead scales with the number of participating agencies. Three agencies sharing data requires roughly nine bilateral trust relationships to maintain, not three.
- Shared infrastructure creates shared liability. When a breach occurs, every agency whose data was in the system will be named in the incident report and the congressional inquiry.
- Standards become anchors. Once fifty counties are conforming to a data model, changing that model for good reasons requires a change management program that can take years, regardless of how technically straightforward the change is.
- Federated models, where agencies retain their own systems but expose standardized query endpoints, often preserve more autonomy and political will than centralized data lakes, even if they are architecturally messier.
The programs that have worked consistently share one characteristic: a named data governancedata governanceData governance is the set of policies, roles, and processes that ensure data is accurate, secure, well-defined, and used responsibly across an organization.View full definition → board with cross-agency membership, decision authority, and a published escalation path. Technology without that structure does not interoperate. It integrates once and then diverges quietly until the next program review finds the records no longer match.
Interoperability between public agencies is a governance problem with a technology expression, not the other way around. CDOs who sequence it correctly, defining meaning before building pipes and securing legal authority before writing code, finish programs. Those who invert the sequence spend the back half of every project trying to retrofit agreements onto infrastructure that was already built without them.
The full course on this sector:Data in Public Sector & Nonprofit.
Frequently asked questions
What is NIEM and do all U.S. federal agencies have to use it?
NIEM (National Information Exchange Model) is a shared data dictionary that defines common terms across justice, health, immigration, and emergency management domains. It is not a universal federal mandate but is widely referenced in cross-agency data exchange programs, particularly in justice and public safety, because it resolves semantic mismatches before technical integration begins.
How long does a cross-agency data-sharing agreement typically take to finalize in the U.S. federal government?
Federal interagency data-sharing agreements commonly take six to eighteen months to execute, depending on the agencies involved and the statutory authorities being invoked. Privacy Act interpretations, agency legal counsel review, and procurement requirements under the Federal Acquisition Regulation all add time that CDOs need to build into program schedules from the start.
What is the difference between semantic interoperability and technical interoperability in government data programs?
Semantic interoperability means two agencies agree on what a data field means, for example how "household" or "incident" is defined, before any data is exchanged. Technical interoperability covers the formats and protocols used to transmit data, such as FHIR for health records. In practice, semantic agreement takes longer and breaks more programs than the technical layer does.
Is a federated data model better than a centralized data lake for cross-agency sharing?
Federated models, where agencies keep their own systems but expose standardized query endpoints, often preserve more political will and agency autonomy than centralized lakes. The tradeoff is architectural complexity and less consistent data quality across sources. For programs where shared liability is a concern, or where agencies have strong competing interests, federation is frequently the more durable choice.
Go deeper
The lessons that take this article further, free to read.
- 1Interoperability and shared standards across agenciesData in the public sector
- 2Mapping the public data landscape: registries, admin records, and survey dataData in the public sector
- 3Governing privacy and equity on legacy systemsData in the public sector
- 4Writing consent and data-sharing agreements that survive an auditData in the public sector
- 5Building open data that citizens and journalists actually useData in the public sector
Sources
- Enterprise AI desperately needs to protect data and models. Here’s how confidential AI could do it.
- Meta hired MongoDB’s CEO to build its enterprise AI business — but Llama is missing
- AI Slop Is Already in Your Training Dataset. I Tested Three Ways to Spot It.
- Fivetran + dbt Labs Announces New Capabilities to Make Enterprise Data Agent-Ready at dbt Summit 2026
- Everything we announced at dbt Summit and why it matters
- We built dbt State to stop rebuilding what hadn't changed
- Celebrating the 2026 dbt partner of the year winners
- Spot New Tech Skills Emerging From the Workforce
- Building on AI’s Unfinished Foundation
- Databricks processes your data. dbt defines what it means
- dbt Core v1.12 is GA
- Model for the token, not the table
- How dbt State cuts warehouse compute and speeds up every run
- dbt Summit 2026: the keynotes and product sessions
Finished reading?
Validate your read to earn XP and feed your radar.