+150 XP

Building a pharma data governance operating model

The hook: two datasets, two very different owners

Picture a data governance council meeting at a mid-size pharma company in 2026. Item one: who owns the HCP (Health Care Professional) engagement data generated when a sales rep logs a conversation with a cardiologist into the CRM (Customer Relationship Management system)? Item two: who owns the biomarker data pulled from a Phase II trial's blood samples, now sitting in a lab information system?

These sound like similar "who owns the data" questions. They are not. The HCP engagement data is jointly claimed by Commercial (which needs it for targeting) and Medical Affairs (which needs it for compliance review of promotional interactions). The biomarker data is claimed by R&D (which needs it for the primary endpoint analysis) and increasingly by Commercial (which wants early biomarker signals for launch planning).

The council's job is not to pick a single "owner" in the IT sense. It is to assign data stewardship (the operational responsibility for data quality, access rules and lifecycle), separate from data ownership (the accountability for how the data is used), and to document both in a policy the whole company can audit against.

Why pharma needs a formal operating model, not just a policy PDF

Pharma sits at the intersection of the strictest privacy regimes and the strictest product regulations. A data governance operating model is the concrete structure (roles, committees, tools, review cadences) that makes policy enforceable day to day.

Three regulatory forces make this non-optional:

  • GDPR (General Data Protection Regulation, EU, effective 2018) governs any personal data tied to EU patients or HCPs, including pseudonymized trial data in many interpretations.
  • HIPAA (Health Insurance Portability and Accountability Act, US, 1996) governs protected health information (PHI) held by covered entities and their business associates, relevant when pharma companies receive real-world data from health systems.
  • 21 CFR Part 11 (US FDA regulation on electronic records and signatures) governs the integrity of data supporting regulatory submissions, meaning trial data governance failures can jeopardize an approval, not just trigger a privacy fine.

Layer onto this the EU's Clinical Trials Regulation (536/2014) and the FDA's data integrity guidance, and you see why "who can touch this dataset, and how do we prove it" is a board-level question, not just an IT ticket.

The core roles: who actually does what

A working operating model needs named roles, not just a policy statement. The standard pharma structure looks like this:

Data Owner: A senior business leader (e.g., VP of Clinical Operations, VP of Commercial Analytics) accountable for a domain's data. Owners approve access policies and sign off on new uses.

Data Steward: An operational role (often embedded in the function, not IT) who manages day-to-day data quality, metadata, and access requests. A clinical data steward ensures trial datasets are clean before a database lock; a commercial data steward ensures HCP records are deduplicated and consent flags are current.

Data Governance Council: A cross-functional body, typically including R&D, Medical, Commercial, Legal, Privacy/Compliance and IT, meeting monthly or quarterly to resolve disputes (like the HCP-vs-biomarker example above) and approve new data-sharing agreements.

Data Protection Officer (DPO): A legally mandated role under GDPR for many pharma companies, independent of operational data owners, responsible for privacy compliance oversight.

Qualified Person / Data Integrity Lead: In GxP (Good Practice, e.g. Good Clinical Practice, Good Manufacturing Practice) contexts, a role accountable for ensuring data supporting regulatory filings meets ALCOA+ principles (Attributable, Legible, Contemporaneous, Original, Accurate, plus Complete, Consistent, Enduring, Available).

Access tiers: not everyone gets the same view

A practical governance model defines access tiers tied to role and purpose, not just seniority. A common structure:

TierWhoExample Access
Tier 1: IdentifiedTrial site staff, treating physiciansFull patient identifiers, needed for safety follow-up
Tier 2: PseudonymizedBiostatisticians, data managersSubject IDs replace names; re-identification possible only via a separate key
Tier 3: Aggregated/De-identifiedCommercial analytics, market accessCohort-level summaries, no individual-level re-identification risk
Tier 4: Public/SyntheticExternal partners, academic collaboratorsSynthetic or fully anonymized datasets

This tiering is what makes a data-sharing agreement (DSA) enforceable. A DSA is a contract specifying what data moves between two parties (say, a pharma company and an academic hospital running a real-world evidence study), under what tier, for what purpose, and for how long. Every DSA should map explicitly to one of these tiers rather than using vague language like "de-identified data" without defining the standard used (HIPAA Safe Harbor and the EMA's anonymization guidance differ in specifics).

A concrete worked example: routing a data request

Say a market access team requests HCP prescribing patterns linked to real-world outcomes data to build a value dossier for payers. The governance workflow:

  1. Request submitted to the relevant Data Steward (Commercial Analytics).
  2. Steward checks: does this require Tier 2 (pseudonymized, needs DPO sign-off) or can Tier 3 (aggregated) satisfy the business need?
  3. If Tier 2 is genuinely required, request escalates to the Data Governance Council with a documented purpose limitation (GDPR Article 5 principle: data collected for one purpose cannot be freely repurposed).
  4. Legal reviews against existing DSAs with the data source (e.g., a claims data vendor like IQVIA or a health system).
  5. Access granted with an expiry date and logged in an access registry, auditable later.

A simple access-control snippet illustrating the logic a governance team might codify in an access-request system:

python
def evaluate_request(purpose, data_tier_requested, requester_role):
    if data_tier_requested == "Tier1_Identified" and requester_role not in ["trial_site_staff", "treating_physician"]:
        return "DENY: identified data restricted to clinical care roles"
    if purpose not in APPROVED_PURPOSES[data_tier_requested]:
        return "ESCALATE: purpose not pre-approved, route to Governance Council"
    return "APPROVE: log access, set 12-month expiry"

This is illustrative logic, not a real compliance engine, but it shows how governance decisions get operationalized into system rules rather than living only in a policy document.

Audits and checks: proving the model works

A governance model is only as good as its audit trail. Practical checks pharma teams run:

  • Consent audits: Sample HCP and patient records quarterly to confirm consent flags match actual data use (critical under GDPR's accountability principle).
  • Access log reviews: Who accessed Tier 1/2 data in the last quarter, and did every access map to an approved purpose?
  • Data lineage checks: For any number appearing in a regulatory submission, can you trace it back to the raw source record (a core Part 11 / ALCOA+ requirement)?
  • DSA expiry sweeps: Are any data-sharing agreements operating past their contractual end date?

The UK's ICO guidance on data protection by design is a good free reference for structuring these audits even outside the UK, since many pharma companies apply one global standard.

Knowledge check

1. In the pharma data governance operating model described, what is the key distinction between data stewardship and data ownership?

2. Why can't the governance council simply assign a single 'owner' to a dataset like HCP engagement data the way IT systems typically assign a system owner?

3. Why is a formal operating model (roles, committees, tools, review cadences) necessary in pharma beyond having a written data governance policy?

MULTIPLE CHOICE

4. Select ALL correct answers about why biomarker data from a Phase II trial creates governance complexity similar to (but distinct from) HCP engagement data.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL correct answers about the regulatory forces that make a formal data governance operating model non-optional for pharma companies.

Select all the correct answers.

Where this breaks down in practice

The most common failure mode is not malice, it is ambiguity. Commercial teams reuse trial biomarker signals for launch targeting without re-checking whether the original patient consent covered secondary commercial use. R&D teams sit on real-world data that Medical Affairs needs for a safety signal review, because no one formally assigned stewardship when the data source was onboarded.

Good governance councils fix this by requiring a data use case register at intake: every new data source gets a documented owner, steward, allowed purposes and tier, before the first analyst touches it, not after a problem surfaces.

Key Takeaways

  • Separate data ownership (accountability for use) from data stewardship (operational quality and access management); pharma disputes usually arise from conflating the two.
  • Build access around defined tiers (identified, pseudonymized, aggregated, synthetic) and require every data-sharing agreement to reference a tier explicitly, not vague terms like "de-identified."
  • Anchor the operating model in named regulations: GDPR and HIPAA for privacy, 21 CFR Part 11 and ALCOA+ for data integrity supporting regulatory submissions.
  • Run recurring audits (consent checks, access log reviews, data lineage traces, DSA expiry sweeps) rather than treating governance as a one-time policy sign-off.
  • Use a cross-functional Data Governance Council to resolve ownership conflicts (R&D vs. Commercial vs. Medical) before they become compliance incidents.

Related articles

Recent articles from the blog that build on this lesson.