+150 XP

Master data and identity resolution: matching HCPs, patients and products

Dr. Elena Marchetti prescribes an anticoagulant to her patients in Milan. In the national prescriber registry she is "MARCHETTI ELENA, Ospedale San Raffaele." In her IQVIA prescription-tracking record (IQVIA is a leading healthcare data and analytics vendor) she appears as "E. Marchetti MD, Cardiology Dept." In the manufacturer's CRM (customer relationship management system used by sales reps), she's "Elena Marchetti, Key Account, Milan Territory 4." Three records. One physician.

If a pharma company's commercial analytics team doesn't know these are the same person, her prescribing volume gets split three ways. She looks like three low-value doctors instead of one high-value specialist. Sales targeting, medical education outreach and compliance monitoring all get it wrong at once.

This is the identity resolution problem, and solving it is the job of master data management (MDM): the discipline of creating one trusted, unique record for each real-world entity (a doctor, a hospital, a drug) across all systems.

Why fragmented identity breaks pharma analytics

Every pharma company runs on data from multiple, independently built sources:

  • Internal CRM (Veeva, Salesforce): sales rep visit logs, sample drops, call notes.
  • Claims and prescription data: from vendors like IQVIA, Symphony Health (part of ICON), or in the US, pharmacy claims switches like Surescripts.
  • Public registries: the NPI (National Provider Identifier, a unique 10-digit ID the US government assigns to every licensed healthcare provider) or, in Europe, national medical council registries that vary by country and often lack a common ID standard.
  • Marketing and medical affairs platforms: conference attendance, speaker bureau payments, medical inquiry logs.

None of these systems were built to talk to each other. Names get misspelled, hospitals merge, doctors change addresses, and abbreviations vary ("St. Mary's Hosp." vs "Saint Mary Hospital"). Without a unifying layer, the same prescriber can generate five to ten fragmented "personas" across a single company's data estate. This is a well-documented issue in pharma commercial operations; industry benchmarking by firms like Reltio and Veeva (both MDM/CRM vendors serving life sciences) routinely cites duplicate HCP (healthcare professional) records as a top data-quality complaint from commercial analytics teams.

The building blocks: HCP/HCO hierarchies

MDM in pharma isn't just deduplication, it's building structured hierarchies:

HCP (Health Care Professional): an individual, like Dr. Marchetti.

HCO (Health Care Organization): the institution she works for, Ospedale San Raffaele, which itself may roll up into a larger hospital network.

A good MDM system links HCP to HCO to network, so a company can answer: "How much total value does the San Raffaele network represent?" not just "How much did this one doctor prescribe?" This matters commercially (accurate account planning) and legally: many transparency and anti-corruption regulations (like the US Sunshine Act, part of the Physician Payments Sunshine Act under the Affordable Care Act, or France's Loi Bertrand / Sunshine française) require companies to report payments to named, correctly identified providers. Misidentification risks compliance failures, not just bad analytics.

Common identifiers used to anchor the match

IdentifierScopeWhat it anchors
NPIUS onlyIndividual provider or organization
GMC/HPD numbersUK, varies elsewhere in EUNational medical registry ID
DEA numberUS onlyControlled substance prescribing authority
IQVIA OneKey IDGlobal, commercial vendorCross-market HCP/HCO master reference

Note there is no single global HCP identifier. This is why commercial MDM vendors (IQVIA OneKey, Veeva Network, Dun & Bradstreet for organizations) exist: they build and license cross-referenced "golden records" that map local IDs to one master ID.

Product identity: NDC and ATC codes

The same fragmentation problem applies to drugs, and it's arguably higher stakes because it affects safety signal detection, not just sales analytics.

NDC (National Drug Code): an 11-digit US FDA identifier, unique to a specific product, strength, and package size. The same molecule from the same manufacturer can have multiple NDCs (different pack sizes, formulations).

ATC (Anatomical Therapeutic Chemical) classification: a World Health Organization system that groups drugs by the organ system they act on and their therapeutic/chemical properties (maintained by the WHO Collaborating Centre for Drug Statistics Methodology). Used globally, especially in Europe, for epidemiological and utilization studies.

A worked example of why this matters: if a pharmacovigilance team (drug safety monitoring) is scanning adverse event reports and one adverse event database logs a product by brand name, another by NDC, and another by generic molecule name, a genuine safety signal can get diluted across three "different" product records instead of surfacing as one clear pattern.

Example: is this the "same" product across systems?

System A (claims):   NDC 00069-0420-30 -> "Zithromax 250mg, 30-count"
System B (EHR):       "Azithromycin 250 MG Oral Tablet"
System C (WHO ATC):   J01FA10 (Azithromycin, anti-infective class)

MDM/product master must resolve all three to one product entity,
while preserving the ability to roll up to molecule (azithromycin)
or therapeutic class (J01FA, macrolides) for epidemiological analysis.

Free reference tool: the WHO ATC/DDD Index lets you look up any ATC code and cross-check classification logic, useful for anyone doing pharma market analysis.

How matching actually works: deterministic vs probabilistic

Two core techniques, often used together:

Deterministic matching: exact-match rules on a trusted identifier. If two records share the same NPI, they're the same entity. Fast, precise, but fails when the identifier is missing or inconsistent (common outside the US).

Probabilistic matching: scoring similarity across multiple weaker signals (name similarity, address proximity, specialty, phone number) when no single reliable ID exists. This is where fuzzy matching algorithms and, increasingly, machine learning classifiers come in, trained to output a confidence score ("87% likely the same HCP") rather than a binary yes/no.

A pragmatic pharma MDM pipeline layers both: try deterministic match first (cheap, high confidence), fall back to probabilistic matching for the remainder, and route low-confidence matches to human data stewards for manual review.

Knowledge check

1. In the Dr. Elena Marchetti example, what is the core business consequence of failing to resolve her identity across systems?

2. What is the primary purpose of master data management (MDM) in the pharma context described?

3. Why is identity resolution especially difficult across pharma data sources like CRM, claims data, and public registries?

MULTIPLE CHOICE

4. Select ALL correct answers about sources of fragmented identity data in pharma analytics mentioned in the lesson.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL correct answers about why the same physician might appear differently across data systems.

Select all the correct answers.

Data quality metrics that matter here

Generic "data quality" is too vague to manage. Pharma MDM teams track specific, auditable metrics:

  • Match rate: percentage of records successfully linked to a golden record. A common industry target cited by MDM vendors is above 90 to 95 percent match rate for core HCP data, though this is a vendor-cited benchmark, not a regulatory standard, and actual rates vary widely by market maturity.
  • Duplicate rate: percentage of "new" records created that are actually duplicates of existing entities. Lower is better; this is the direct inverse symptom of the Dr. Marchetti problem.
  • Survivorship accuracy: when merging duplicate records, which field values "win" (most recent address? most complete specialty tag?). Poorly governed survivorship rules silently corrupt the golden record.
  • Data freshness / decay rate: HCP data decays fast. Providers change addresses, affiliations, and specialties. Industry commentary from MDM vendors often cites that a meaningful share of HCP contact data becomes outdated within a year, a useful directional estimate, though exact figures vary by source and market.
  • Golden record completeness: percentage of mandatory fields (specialty, primary affiliation, license status) populated and validated.

These metrics should appear on a data governance dashboard reviewed regularly by a data governance function, often reporting into a Chief Data Officer or equivalent, distinct from IT and distinct from commercial operations, precisely so that no single business unit's incentives distort the golden record.

🎬 [VIDEO: "What is Master Data Management?" - youtube.com - search for this title from IBM Technology's channel for a vendor-neutral, plain-language walkthrough of MDM concepts applicable beyond pharma]

Key Takeaways

  • Master data management resolves fragmented identifiers (NPI, IQVIA OneKey ID, internal CRM IDs) into one "golden record" per HCP, HCO or product, preventing split or duplicated commercial and safety analysis.
  • HCP/HCO hierarchies matter for both commercial account planning and compliance reporting (e.g., US Sunshine Act payment disclosures), which require correctly attributing payments to named individuals and institutions.
  • NDC codes identify specific US drug products and packages; ATC codes classify drugs therapeutically worldwide (maintained by WHO). Product identity fragmentation can dilute pharmacovigilance signal detection, not just distort sales reporting.
  • Matching combines deterministic methods (exact ID match) with probabilistic methods (fuzzy, confidence-scored matching) plus human data steward review for low-confidence cases.
  • Track concrete data quality metrics (match rate, duplicate rate, survivorship accuracy, data freshness) under independent data governance, not left to commercial teams whose incentives may favor volume over accuracy.