+150 XP

Mapping the pharma data landscape: sources, vendors and standards

A single new drug launch in the United States can require a pharma company to license data from six or more separate vendors, each charging six or seven figures annually, just to answer the question "is this drug being prescribed the way we intended?" That fragmented, expensive, overlapping data supply chain is the real infrastructure of the pharmaceutical industry, and almost nobody outside the industry has mapped it. This lesson does.

Why the data landscape is fragmented by design

Pharma data is scattered across the supply chain because no single entity sees the whole patient journey. A pharmacy sees the fill. A payer sees the claim. A hospital sees the diagnosis. A lab sees the biomarker. Privacy law (in the US, HIPAA, the Health Insurance Portability and Accountability Act of 1996) keeps these systems from freely merging identifiable data, so a commercial layer of vendors emerged to aggregate, de-identify, and resell fragments of the picture.

Understanding pharma "data fluency" starts with knowing who owns which fragment.

The core data sources, mapped

Prescription and pharmacy data

  • IQVIA: the largest health data and analytics vendor globally, formed from the 2016 merger of IMS Health and Quintiles. IQVIA aggregates prescription (Rx) data from roughly 90%+ of US retail pharmacies (estimate, vendor-reported) via its National Prescription Audit and Xponent products. It licenses this at the brick-and-mortar, prescriber, and payer level.
  • Symphony Health (now part of ICON plc): IQVIA's main historical competitor for longitudinal patient-level Rx and medical claims data, often used as a cross-check because methodology differs.

Where the blind spot is: cash-pay prescriptions, specialty pharmacy dispensing outside standard networks, and drugs distributed via hospital "buy and bill" channels (common for infused oncology drugs) are undercounted in standard Rx audits.

Medical and pharmacy claims

Claims data comes from payers (insurers, pharmacy benefit managers) and clearinghouses. Key commercial aggregators include IQVIA, Komodo Health, and Definitive Healthcare. Claims capture diagnosis codes (ICD-10, International Classification of Diseases, 10th revision), procedure codes (CPT, Current Procedural Terminology), and payment data, but only for insured populations, and only after billing, which lags real-world treatment by weeks or months.

Electronic health records (EHR)

EHR data comes from clinical systems like Epic and Oracle Health (formerly Cerner). Companies such as Komodo, Truveta (a consortium owned by US health systems), TriNetX, and Optum aggregate de-identified EHR data for research and commercial use. EHR data adds lab values, clinical notes, and vitals that claims data lacks, but coverage depends entirely on which health systems contribute, creating geographic and demographic skew.

Clinical trial infrastructure: CTMS and EDC

A CTMS (Clinical Trial Management System, e.g., Veeva Vault CTMS, Oracle Siebel CTMS) tracks trial operations: site activation, patient enrollment, monitoring visits. An EDC (Electronic Data Capture) system (e.g., Medidata Rave) captures the actual patient data collected during a trial. These are the systems of record for regulatory submissions and are subject to strict data integrity rules under 21 CFR Part 11 (the FDA regulation governing electronic records and signatures).

Lab and manufacturing data: LIMS

A LIMS (Laboratory Information Management System) tracks samples, assays, and quality control results in both R&D labs and manufacturing quality control. LIMS data feeds regulatory batch records and is central to GMP (Good Manufacturing Practice) compliance. Vendors include LabVantage, Thermo Fisher's SampleManager, and LabWare.

Post-market safety: FAERS

FAERS (FDA Adverse Event Reporting System) is a free, public database of adverse event reports submitted by manufacturers, healthcare providers, and consumers. It is the backbone of pharmacovigilance (the science of drug safety monitoring) but is voluntary for consumer reports and subject to major underreporting and reporting bias. Explore it directly here: FDA FAERS Public Dashboard.

Europe's equivalent is EudraVigilance, run by the European Medicines Agency (EMA), covering the European Economic Area.

A quick worked example: why blind spots cost money

Say a brand team wants total US patient count on their drug. They license IQVIA Xponent (retail Rx) showing 100,000 unique patients filled a prescription in Q1.

But their drug is also dispensed via specialty pharmacy and hospital infusion, not fully captured in standard retail Rx audits. A claims-based cross-check (via Komodo or Symphony) shows 118,000 unique patients with a paid medical or pharmacy claim in the same window.

The gap: 18,000 patients, or 15% of true volume, invisible in the first dataset alone. That's the concrete cost of relying on a single vendor: mis-sized market estimates feeding directly into forecasting and sales compensation decisions.

Retail Rx count (IQVIA Xponent):        100,000
Claims-based count (Komodo/Symphony):   118,000
Undercount gap:                          18,000
Undercount rate: 18,000 / 118,000 ≈ 15.3%

Data quality and governance metrics that matter

Once you have the data, you need to know if you can trust it. Core metrics pharma data teams track:

  • Completeness: % of expected fields populated (e.g., % of claims records with a valid NDC, National Drug Code).
  • Timeliness/lag: days between event and data availability. Claims data often lags 60 to 90 days (estimate); EHR data can lag weeks; FAERS reports can lag months after an adverse event.
  • Match rate: % of patient records successfully linked across datasets via tokenization (a privacy-preserving hashing method used by vendors like Datavant to link de-identified records without exposing identity).
  • Concordance: agreement rate between two independent sources measuring the same thing (e.g., IQVoia vs. Symphony patient counts for the same drug, as in the worked example above).
  • Audit trail integrity: for CTMS/EDC data, whether every change is logged with a timestamp and user ID, a hard requirement under 21 CFR Part 11.

Knowledge check

1. Why did a commercial layer of data vendors emerge in the pharma industry rather than a single entity tracking the full patient journey?

2. A pharma brand team wants to know if their infused oncology drug is being appropriately prescribed, but standard Rx audit data shows surprisingly low volume. What is the most likely explanation?

3. Why might a pharma analytics team license data from both IQVIA and Symphony Health for the same market instead of just one?

MULTIPLE CHOICE

4. Select ALL correct answers about why the pharma data landscape is described as 'fragmented by design.'

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL correct answers about known blind spots or limitations in standard prescription (Rx) audit data.

Select all the correct answers.

Standards that hold the ecosystem together

Without shared standards, none of this data would be interoperable:

  • HL7 FHIR (Fast Healthcare Interoperability Resources): the modern data exchange standard for clinical data, increasingly mandated by US interoperability rules from the Office of the National Coordinator for Health IT (ONC).
  • CDISC standards (SDTM, ADaM): mandatory formats for submitting clinical trial data to the FDA and EMA.
  • NDC and RxNorm: drug identification codes used across claims and Rx data to identify exactly which product and package were dispensed.
  • ICD-10 / ICD-11: diagnosis coding, maintained by the World Health Organization, used globally though the US uses a modified clinical version (ICD-10-CM).
  • MedDRA (Medical Dictionary for Regulatory Activities): the standard terminology for coding adverse events in FAERS and EudraVigilance.

A useful free primer on how these regulatory data standards fit together: CDISC Standards Overview.

Licensing costs and the vendor power balance

IQVIA's scale gives it significant pricing power. Large pharma companies commonly maintain multi-million dollar annual contracts (estimate, industry-reported, exact figures are confidential) covering Rx data, claims, and analytics tools. Mid-size and smaller biotechs often cannot afford full IQVIA suites and instead combine cheaper point solutions (Symphony/ICON, Komodo, specialty claims vendors) with more manual reconciliation, a real competitive disadvantage in commercial analytics maturity.

This concentration also creates a single point of methodological risk: if most of the industry benchmarks against one vendor's numbers, an error or a definitional change in that vendor's methodology ripples across sales forecasts, market share estimates, and even executive compensation formulas tied to "market share versus IQVIA."

🎬 [VIDEO: "How IQVIA Uses Data to Improve Healthcare" - youtube.com - search this title on YouTube; IQVIA's own explainer on how it aggregates and licenses pharma data, useful for seeing vendor positioning firsthand]

Key Takeaways

  • Pharma data is fragmented across pharmacy, claims, EHR, clinical trial, lab, and safety systems by regulatory and structural design (HIPAA-driven de-identification, siloed care delivery), not accident.
  • IQVIA and Symphony Health/ICON dominate Rx and claims data licensing; no single vendor sees the full patient journey, so cross-checking sources is standard practice, not paranoia.
  • Specialty pharmacy, cash-pay, and hospital buy-and-bill channels are recurring blind spots in standard retail Rx audits, and can undercount true patient volume by double digits percentage-wise.
  • Data quality is measured concretely: completeness, timeliness/lag, cross-vendor concordance, and match rate via tokenized linkage (e.g., Datavant).
  • Standards like HL7 FHIR, CDISC, NDC/RxNorm, and MedDRA are what make this fragmented vendor landscape interoperable at all, and are non-negotiable for regulatory submissions.

Related articles

Recent articles from the blog that build on this lesson.