# Measuring the ROIROIReturn on Investment: the ratio of net profit to the cost of an investment. A 300% ROI means each dollar invested returns $3.View full definition → of clean data on loss ratios
A single mistyped digit in a policyholder's garage address can shift a vehicle from a low-theft suburb to a high-theft zip code, quietly under-pricing a policy for years. Multiply that error across a book of 500,000 policies and you have a measurable, budget-line problem, not an IT nuisance. This lesson builds the quantitative case for treating data qualitydata qualityThe degree to which data is fit for purpose: accurate, complete, consistent, timely, valid and unique. Poor quality data undermines analytics, reporting and AI.View full definition → as a loss-ratio lever, not a compliance chore.
The loss ratio (claims paid divided by premiums earned) is the scoreboard, but data qualitydata qualityThe degree to which data is fit for purpose: accurate, complete, consistent, timely, valid and unique. Poor quality data undermines analytics, reporting and AI.View full definition → is upstream of it in two specific ways:
Both are measurable with data-quality metrics before they ever surface in the loss ratio. That lead time is the whole ROIROIReturn on Investment: the ratio of net profit to the cost of an investment. A 300% ROI means each dollar invested returns $3.View full definition → argument: fix the data, see the ratio move months later, and you can trace the causal chain.
Before measuring quality, know the sources being scored:
Each of these has its own error rate, and errors compound as data moves from intake to rating to claims.
Insurers should track these at the field level, not just "overall data qualitydata qualityThe degree to which data is fit for purpose: accurate, complete, consistent, timely, valid and unique. Poor quality data undermines analytics, reporting and AI.View full definition →":
| Metric | Definition | Typical target (estimate, 2025-26) |
|---|---|---|
| Address match rate | % of policy addresses that geocode to a verified postal/geospatial reference | 95%+ |
| VIN decode accuracy | % of vehicle records where VIN decodes to a valid make/model/year/trim | 97%+ |
| Field completeness | % of mandatory rating fields populated at bind | 98%+ |
| Duplicate policy/claimant rate | % of records flagged as duplicates | under 1% |
| Data latency | time between real-world change (e.g., address change) and system update | under 30 days |
These are leading indicators. A drop in address match rate today predicts mis-rated renewals in 60 to 90 days, well before the loss ratio reports it.
Here is the calculation leadership actually wants to see.
Assumptions (illustrative, not sector-specific claims):
Step 1: Policies affected today
500,000 × 8% = 40,000 mis-rated policies
Step 2: Premium leakage today
40,000 × $1,400 × 12% = $6.72 million in annual under-collected premium
Step 3: Apply a 5-percentage-point accuracy improvement (error rate falls from 8% to 3%)
New mis-rated count: 500,000 × 3% = 15,000
Reduction: 25,000 fewer mis-rated policies
Step 4: Recovered premium
25,000 × $1,400 × 12% = $4.2 million per year in recovered rating accuracy
That $4.2 million doesn't appear as "revenue." It appears as a lower loss ratio, because claims costs on those 25,000 policies no longer run against an underpriced premium base. If the data-quality program (matching software, address verification APIs, VIN decode services) costs $600,000 to $900,000 a year, the payback is well under a quarter, even before counting claims-leakage savings.
Address and vehicle errors also inflate claims leakage, defined as claim payments above what a correctly handled claim would cost. Industry estimates commonly cited by consultancies like McKinsey put avoidable claims leakage at 5% to 10% of incurred losses for P&C (property and casualty) insurers, driven partly by bad data: wrong vehicle valuation, incorrect policy limits pulled at FNOL, missed prior-claim history.
Simple leakage calculation:
If incurred losses are $300 million and leakage is conservatively 6%, that's $18 million. If a third of leakage traces to data errors (a reasonable, commonly cited split for identifiable root causes), cleaning up VIN decode and claimant-matching accuracy targets roughly $6 million annually. This is the number to pair with the premium-recovery figure above when building a board case.
A one-time data cleanse decays. Governance metrics keep it from decaying:
In the US, the National Association of Insurance Commissioners (NAIC) increasingly scrutinizes data governancedata governanceData governance is the set of policies, roles, and processes that ensure data is accurate, secure, well-defined, and used responsibly across an organization.View full definition → during market conduct exams. In Europe, Solvency II's Article 19 explicitly requires insurers to demonstrate data "appropriateness, completeness, and accuracy" for reserving and capital models, non-compliance can trigger capital add-ons, a direct financial consequence of poor governance.
Knowledge check
1. Why does a mis-typed garage address on a policy create a measurable financial problem rather than just a minor data entry issue?
2. What is the key distinction between 'mis-rating' and 'claims leakage' as two pathways by which data quality affects the loss ratio?
3. The lesson argues that data-quality metrics can be measured 'months before' the loss ratio moves. Why is this lead time central to the ROI argument for clean data?
4. Select ALL correct answers describing ways that poor data quality can cause 'claims leakage.'
Select all the correct answers.
5. Select ALL correct answers about the categories of datasets identified as relevant to measuring data quality's impact on the loss ratio.
Select all the correct answers.
To make the ROIROIReturn on Investment: the ratio of net profit to the cost of an investment. A 300% ROI means each dollar invested returns $3.View full definition → case durable, pair quality metrics with outcome metrics on a single dashboard reviewed monthly:
Leading indicators Lagging indicators
------------------ -------------------
Address match rate -> Loss ratio by territory
VIN decode accuracy -> Vehicle-class loss ratio
Field completeness -> Renewal mis-rate rate
Duplicate claimant rate -> Claims leakage %
Data latency -> Subrogation recovery rateThe point isn't a fancy tool. It's the causal pairing: every quality metric on the left should have a named, trackable outcome on the right, so leadership can see the lag between cleaning the data and the loss ratio responding.