One HCP, six records: why identity resolution is pharma's most expensive data problem

A single cardiologist can exist as six different entities across a pharma company's CRM, claims data, and prescriber analytics systems, and none of them match. Until identity resolution works in practice, every downstream decision, from sampling allocations to pharmacovigilance reporting, is built on a fractured foundation.

Listen to the podcast

4 min

Chapters

Key takeaways

  • Make the federal NPI registry the source of truth for prescriber identity and force every other system to point at it.
  • Run a match against the federal file nightly instead of treating deduplication as a one-off project that decays within 18 months.
  • Write down explicit survivorship rules so you know which record wins when two versions disagree on address or name.
  • Name one person accountable for the prescriber master record, with budget, before buying any matching tool.
  • Pull every record for your top 10 prescribers by revenue by hand and count the duplicates to get the dollar figure your board needs.
Read the full transcript

Host:This is Leader's Insights. On the table, 1 HCP, 6 records. Why identity resolution is pharma's most expensive data problem. 6 records for 1 cardiologist. That's the number that stopped me.

Expert:6 is the average I see when I audit a mid-sized pharma's systems. I've seen 11. 1 cardiologist in Ohio spelled 4 ways, 2 license numbers, 3 addresses because she moved practices twice and a middle initial that shows up in the claims feed but not the CRM, the customer system where reps log their visits.

Host:And none of them talk to each other.

Expert:None. They can't. The CRM knows Dr. J. Ramirez, 400 Elm. The claims data, the record of what she actually prescribed, bought from a data broker, knows Jennifer Ramirez, MD, NPI, and the NPI. The NPI is the National Provider Identifier, the federal 10-digit number every prescriber gets. And the Analytics Warehouse has a third version built by a vendor 5 years ago that nobody remembers configuring.

Host:So what breaks first?

Expert:Sampling. You send free drug samples based on who's a high prescriber. If your cardiologist is split across 6 records, each one looks like a low prescriber. So you under-sample your best customer and over-sample 4 ghosts. I watched a company waste roughly a third of a sample budget that way. That's real product, temperature controlled, expiring on shelves.

Host:A third. Is that your number or someone's marketing deck?

Expert:That one's mine, from the audit. But there's a related figure from DBT Labs. They sell data tooling, so read it with one eyebrow up. Claiming companies spend a huge share of analyst time just reconciling identities before any actual analysis. Directionally, I believe it. Cross-check against MIT Sloan Management Review, which has written for years that data quality, not data volume, is where the money leaks.

Host:Let me push on the scary part. You said pharmacovigilance.

Expert:Pharmacovigilance. Safety monitoring, tracking adverse events after a drug ships. If a patient has a bad reaction and the report ties to record number 3, but the prescriber's actual history lives in record 1, you might miss a pattern. Regulators don't accept our systems didn't match. As an excuse.

Host:Has anyone actually been burned by that on the record?

Expert:Nobody puts our identity resolution failed in a press release. What you see is the downstream symptom. A delayed signal. A warning letter about incomplete reporting. I've sat in the room after one. The root cause traced back to duplicate prescriber records nobody had merged. So the lesson the good operators already learned isn't by a matching tool. Its identity is infrastructure. Treat it like plumbing.

Host:Explain the plumbing thing because every vendor says by their tool.

Expert:A tool matches records once. Plumbing means you decide permanently which system is the source of truth for a prescriber's identity, usually the NPI registry because it's federal and stable. And every other system bends to it. Novartis and a few others move to what people call a master data approach years ago. One golden record. Everything else points at it. Boring. Unsexy. Works.

Host:But NPI changes. People move, retire, change specialties.

Expert:Right. So you don't do it once. You run a match every night against the federal file, and you keep a survivorship rule. Wania, plain terms. When two records disagree on her address, which one wins? You write that rule down. Most companies never do. They match once in a big project. Declare victory. And 18 months later, they're back to six records because the plumbing had no maintenance.

Host:So the failure isn't technical. It's that nobody owns it.

Expert:Exactly. Identity has no natural owner. Sales wants their CRM. Analytics wants their warehouse. Safety wants theirs. Everybody assumes someone else deduplicated. I tell every board, name one person accountable for the prescriber master record with budget. Or don't bother buying anything. Give me the one thing a listener does Monday

Host:morning. Pick your top 10 prescribers by revenue. Pull

Expert:every record for each across your systems by hand. Count them. If any of those 10 shows up more than once, and they will, you now have the exact dollar figure your board needs to fund the fix. Don't theorize the problem.

Host:Photograph it. Photograph it. That's the takeaway. What we read for this one, DBT Labs vendor data tooling MIT Sloan Management Review. We'll stop there. The whole CDO track and order is waiting at MBA-training.com.

Identity resolution in pharma data systems refers to the process of determining that two or more records across separate data sources refer to the same real-world entity: a specific healthcare professional, a patient, or a product. The concept sounds straightforward. The execution is not, and the consequences of failure run from wasted commercial spend to regulatory breach.

Which identity domains must a pharma CDO resolve?

The pharmaceutical industry maintains three distinct identity domains simultaneously, and each one carries its own stakes.

For healthcare professionals, the commercial stakes are direct. A medical science liaison visiting "Dr. Sarah Chen" at Brigham and Women's Hospital should draw on the same profile as the CRM record from Veeva, the prescriber data purchased from IQVIA, the speaker bureau record from a third-party compliance vendor, and the adverse event report submitted to FDA. If those records don't resolve to one entity, the company might over-promote to the same physician across three channels (an anti-kickback exposure), under-sample a high-decile prescriber because her records are split across two territories, or, in a pharmacovigilance context, miss a safety signal because duplicate reports are treated as separate cases.

For patients, the stakes shift to safety and privacy together. Longitudinal patient data, drawn from specialty pharmacy dispensing records, HEOR datasets like IQVIA's Real World Solutions or Optum's claims data, and patient support program touchpoints, only yields valid outcomes evidence if the same patient's records are linked across time. Fragmented identity means fragmented therapy timelines, which corrupts both real-world evidence studies submitted to regulators and the internal analytics that guide patient support program design. Misidentification in a pharmacovigilance report can mean a duplicate adverse event inflates a drug's safety signal. The inverse, merging two patients who are distinct, can suppress one.

For products, the problem is less intuitive but equally costly. A single biologic may carry an NDC code in the US dispensing channel, a different identifier in the 340B covered entity system, a CIN from a GPO catalog, and an entirely separate product code in the EU under EMA's IDMP (Identification of Medicinal Products) framework. Failure to resolve these to a single product identity means government price reporting, specifically ASP and AMP calculations for Medicaid and Medicare, draws from mismatched product records. The financial exposure from that failure is not theoretical: CMS audits have resulted in significant manufacturer repayments, and false claims liability follows.

How do pharma teams match HCP records across sources?

The mechanics fall into three broad approaches: deterministic matching, probabilistic matching, and graph-based resolution.

Deterministic matching is exact: two records that share an NPI number (for US HCPs) or a DEA number refer to the same entity. Full stop. NPPES, the National Plan and Provider Enumeration System, is the canonical US source. Pharma data teams that build their HCP master on NPI as the primary key and enrich from there, via OneKey from IQVIA or DDD from Wolters Kluwer Health, are working with the cleanest starting point available. The limit is obvious: NPI only covers US prescribers, does not exist for key opinion leaders who don't prescribe (lab researchers, KOL-type PhDs), and provides no cross-border linkage for global companies tracking HCPs across the EU.

Probabilistic matching activates when deterministic keys are absent or conflicting. The algorithm scores record pairs on weighted attributes: surname, given name, specialty, address, phone, graduation year, affiliated institution. A Jaro-Winkler string distance on "Catherine Moreau" versus "C. Moreau-Dubois" at the same hospital address in Lyon might return a match score of 0.87. The matching engine applies a threshold, say 0.80, to declare a match, a lower threshold to flag for human review, and below that to treat records as distinct.

A concrete example: a European pharma company launches a new oncology biologic across France, Germany, and the Netherlands. Their medical affairs team has 14,000 HCP records from three national distributors, a conference attendance list from ESMO 2025, and a speaker engagement database. The same oncologist, Dr. Moreau, appears four times across those sources, spelled differently in each, affiliated with two hospitals (she holds a joint appointment), and with a French RPPS identifier in one record and an email address in another. A deterministic pass matches two of the four records via RPPS. The probabilistic layer resolves the third via name-address similarity. The fourth, from the ESMO list, carries only her name and an institutional email from her secondary affiliation: it goes to a human review queue. That workflow, deterministic first, probabilistic second, human adjudication third, is the production standard at companies like Roche and Sanofi for HCP master data maintenance.

Graph-based resolution, increasingly adopted as of 2026, models entities and their relationships as nodes and edges. Rather than comparing record pairs in isolation, it asks: does this unresolved HCP record share an institutional affiliation, a co-author relationship, and a prescribing ZIP code cluster with a known entity in the graph? Companies running identity resolution on graph databases (Neo4j is common in this space) find thatthe golden record emerges from relationship evidence, not just attribute similarity, which substantially improves recall for high-mobility HCPs like academic physicians who rotate across institutions.

Accuracy thresholds for pharmacovigilance, commercial targeting and RWE

Identity resolution is not uniformly worth solving to the highest fidelity across every use case.

For pharmacovigilance and ICSR (Individual Case Safety Report) management, there is no acceptable error rate. FDA regulations under 21 CFR Part 314 and EMA's GVP Module VI require accurate patient and reporter identification. A duplicate that doubles an adverse event count for a compound can trigger a regulatory safety review. Invest in deterministic resolution, mandatory human review queues, and a well-governed survivorship policy here without compromise.

For commercial targeting, probabilistic matching to 85-90% accuracy is usually good enough. A 5% error rate in a Veeva CRM HCP universe means some reps visit the wrong decile tier or miss a newly important prescriber. That matters commercially, but it does not create legal exposure by itself, unlessthe errors feed into speaker bureau selection or grant decisions, where anti-kickback scrutiny applies.

For real-world evidence studies, the threshold depends on the regulatory pathway. An RWE study supporting a label extension submitted to FDA under 21st Century Cures Act provisions will face methodological scrutiny on patient matching quality. An internal HEOR analysis informing payer contracting strategy has more latitude.

The honest tradeoff is this: enterprise-grade identity resolution requires sustained investment in data stewardship, not just tooling. Vendors like IQVIA, Veeva, and Komodo Health sell pre-resolved HCP and patient universes as products, which reduces build cost but transfers governance risk. When their matching is wrong, correcting it inside your own systems without overwriting the vendor's primary key is a non-trivial architectural problem.

The CDOs who get this right treat identity resolution as infrastructure owned by data governance, not a feature owned by a single commercial application. That distinction determines whether the golden record holds across a product launch, a merger, and a regulatory inspection.

The full course on this sector:Data in Pharmaceuticals.

Frequently asked questions

What is identity resolution in pharma data?

Identity resolution in pharma is the process of determining that records held in separate systems refer to the same real-world HCP, patient or product. A single oncologist can appear in a Veeva CRM record, IQVIA prescriber data, a speaker bureau file and an adverse event report, and those four records must resolve to one entity before any of the analytics built on them hold.

Is NPI enough to build an HCP master file?

NPI works as a primary key for US prescribers and NPPES is the canonical source, but it stops there. It does not cover key opinion leaders who do not prescribe, such as lab researchers and PhD KOLs, and it gives no cross-border linkage for companies tracking HCPs across the EU, where identifiers like the French RPPS apply instead.

What happens when product identities do not match in price reporting?

Mismatched product records feed wrong inputs into ASP and AMP calculations for Medicaid and Medicare. A biologic can carry an NDC code in the US dispensing channel, another identifier in the 340B system, a GPO CIN and a separate EU code under EMA's IDMP framework. CMS audits over these gaps have led to significant manufacturer repayments and false claims liability.

Should pharma buy pre-resolved HCP data or build it in-house?

Buying pre-resolved universes from IQVIA, Veeva or Komodo Health cuts build cost but transfers governance risk to the vendor. When the vendor's matching is wrong, correcting it inside your own systems without overwriting their primary key becomes a hard architectural problem, which is why identity resolution belongs to data governance rather than a single commercial application.

Go deeper

The lessons that take this article further, free to read.

  1. 1Master Data Management in practice: styles, tools, and the Golden RecordData governance & compliance
  2. 2Data quality metrics that matter: completeness, latency and lineageData in pharma
  3. 3Mapping the pharma data landscape: sources, vendors and standardsData in pharma
  4. 4Consent, de-identification and the limits of anonymous dataData in pharma
  5. 5Marketing on a leash: promotional compliance and the anti-kickback minefieldPharma: how the sector works

Finished reading?

Validate your read to earn XP and feed your radar.