Substantiation files and the evidence behind every claim
A reviewer only has to ask one question about the line "40% fewer revisions": show me the page. Not the study title, the page. The table, the arm labels, the follow-up window, the confidence interval. If that page takes more than a minute to produce, the claim is already fragile. If the page says something narrower than the brochure, the campaign stops, and every asset built on that sentence stops with it.
What a substantiation file actually is
A substantiation file (also called a claims dossier or evidence file) is the organised bundle of documentary proof behind every promotional statement your company makes. One claim, one traceable source, stored where a reviewer can reachreachThe number of unique people exposed to your message in a given period. Unlike impressions, reach counts each person once, no matter how often they see it.View full definition → it without asking a colleague where it went.
Two rules govern the whole discipline:
- Prior substantiation: the evidence exists *before* the claim is used. Assembling it after a challenge is itself the finding.
- Claim-to-evidence match: the evidence supports the *exact* wording, population, comparator and timeframe of the claim.
"Clinicians tell us" and "everyone says it" are not references. The prohibitions that make these two rules enforceable sit with the authorities and codes the foundations lesson maps; this lesson stays inside the folder.
Grading the reference, not just finding it
Most broken claims are not unsupported. They are supported by something weaker than the sentence implies. Grade every reference before you attach it.
| Tier | What it is | What it can carry |
|---|---|---|
| 1 | Peer-reviewed publication, pre-specified primary endpoint, population matching the indication | The claim as worded |
| 2 | Peer-reviewed, but secondary endpoint, post-hoc subgroup, single-arm or registry | Narrow, hedged wording only |
| 3 | Congress abstract or poster | Provisional wording, and only until the full paper lands |
| 4 | Data on file: internal report, validation study, unpublished analysis | Claims you can defend on request, with the report in hand |
| 5 | Review article, KOL slide deck, another company's brochure | Nothing |
Tier 3 is where files rot quietly. An abstract runs a few hundred words with no full method section, and numbers move between the poster and the publication. If your brochure quotes the poster figure and the paper later reports something smaller, you are circulating printed material that contradicts its own primary source.
Tier 5 is the more frequent error. Someone cites a review that cites the study. The review reads well, so nobody rechecks. When the primary paper is corrected, superseded or retracted, the review still reads well, and the claim keeps running.
Babylon Health showed what a mis-graded reference costs. Its 2018 launch event presented a score of roughly 80% on questions resembling the MRCGP exam as evidence that its triage system worked at the level of a GP. The underlying work was an internal evaluation, not independently validated, and clinicians writing in The Lancet dismantled it on exactly that ground: small question set, no external test data, results publicised before scrutiny. A Tier-4 reference had been asked to carry a Tier-1 sentence, on stage, with journalists in the room.
The three evidence types behind a claim
1. Clinical data
Evidence from studies in humans: randomised trials, registries, post-market clinical follow-up.
"Reduces post-surgical infection rate" is a clinical outcome. It needs a clinical study, ideally comparative, in the population the product is actually indicated for.
2. Bench data
Laboratory or engineering testing, no patients. Tensile strength, flow rates, battery life, sterility, wear simulation.
"Withstands 10 million wear cycles" is a bench claim. It needs a validated test protocol and the raw results, not a trial.
Never let a bench result stand in for a clinical benefit. "Survives 10 million cycles in vitro" does not license "lasts 15 years in patients." That leap is one of the oldest misleading-claim findings on record.
3. Comparative data
Any claim naming a competitor or a category ("versus the leading brand", "the only device that", "40% fewer revisions than standard care"). Highest risk, because it attracts both regulatory attention and a rival's complaint.
These need head-to-head evidence, or an indirect comparison that survives scrutiny: same endpoint definition, comparable populations, comparable follow-up. Your best-case result against a competitor's worst-case is a guaranteed loss.
Tracing a claim back to its page
Work backwards, the way a reviewer will.
| Question | What the file must contain |
|---|---|
| 40% versus what? | Named comparator or defined standard of care |
| In whom? | Population matching the indication |
| Over what period? | Follow-up duration (2 years? 10 years?) |
| Statistically real? | p-value or confidence interval, sample size |
| Same measurement? | Identical definition of "revision" in both arms |
If the study measured revisions at 24 months and the brochure implies lifetime, the claim is broken. If "revision" was defined differently in the two arms, broken. If the 40% has no interval and the sample was 30 patients, probably broken.
A defensible version reads: "In a 24-month randomised study (n=420), revision rate was 40% lower versus [named comparator] (95% CI, p<0.01)." Less punchy, far more survivable.
The harder failure is when the evidence is real and the claim still overshoots. Grail's Galleri blood test detects a cancer signal, and the PATHFINDER study supports that: about 1% of the 6,662 participants had a signal detected, and fewer than half of those were confirmed to have cancer. What no completed randomised trial has yet shown is that screening with it reduces cancer mortality, which is why the NHS-Galleri trial of around 140,000 people in England exists. The file supports "detects a shared cancer signal across many cancer types". It does not support "catches cancer early enough to save your life". Same dossier, two sentences, one of them unfunded.
The claims matrix and where the references live
| Claim ID | Approved wording | Type | Source doc | Expiry / review | Approved by |
|---|---|---|---|---|---|
| C-014 | "40% fewer revisions at 24 months vs. X" | Comparative + clinical | Study RCT-2024-07, Table 2, p.6 | Review Q3 2026 | Med, Reg, Legal |
| C-021 | "Withstands 10M wear cycles" | Bench | Test WR-118 | Static | Reg, R&D |
What makes the matrix work is boring discipline, not the spreadsheet:
- Nothing enters an asset without a Claim ID.
- The source column names the page, table or figure, not just the paper. Reviewers who have to hunt start guessing.
- Every reference is stored as a locked copy in one repository, with the annotated version showing which passage supports which claim. Never a link to a journal site; access changes, paywalls move, and PDFs get replaced silently.
- Retention is a legal requirement, not housekeeping. In UK pharma, certified material and its certificate must be kept for at least three years after final use. In the US, prescription drug promotional pieces go to the FDA at the time of first use on Form 2253, meaning the agency already holds a copy of the thing you would rather forget.
- Claims expire for reasons that have nothing to do with your data. A comparator's label change, a new guideline, a competitor's fresh trial: your 2019 study is unchanged and your claim is now stale.
🎬 [VIDEO: "How the FDA Regulates Medical Device Advertising" - youtube.com - a plain-language walkthrough of promotional rules and the false-or-misleading standard for devices]
The reference pack that travels with the asset
Each approved asset ships with its own pack: the claim list, the graded references, the annotated pages. The sign-off sequence and who owns which signature belong to the pre-launch review lesson. What belongs here is what happens to the pack afterwards, because that is where files break.
- Translation drift. "40% fewer revisions at 24 months" becomes a superlative in a local adaptation, and the affiliate has the parent pack but not the parent wording.
- Format compression. A banner ad has no room for the qualifier, so the qualifier lives one click away. The file needs to show the qualified version as the reader actually encounters it, screenshot included.
- Field decks. A rep adds one slide with a favourite graph from a congress poster. There is no Claim ID, no pack, and it is now the most-used asset in the region.
- Reference chain breaks. The source document is revised and re-issued under a new number, and forty claims still point at the old one.
Audience changes what the evidence has to carry
The test is whether the intended reader is misled, and that shifts with the reader. A dense hazard-ratio claim aimed at interventional cardiologists can be fine. The same claim on a consumer page for an at-home monitor may mislead a lay reader, so it needs plainer wording and often more evidence, not less. The fairness and transparency duties owed to those audiences are covered in the fair-treatment lesson; the point for the file is that one dossier can support a specialist sentence and fail a consumer one.
Four traps recur:
- Implied claims. A photograph of a patient running a marathon makes a mobility claim without words. The image needs its own reference.
- Superlatives. "Best", "most advanced", "gold standard." Each needs proof or gets cut.
- Cherry-picked endpoints. Reporting the endpoint that reached significance and omitting the ones that did not.
- Surrogates dressed as outcomes. A biomarker moved. Nobody lived longer. Say which one you measured.
Knowledge check
1. A marketing team assembles supporting studies only after a regulator challenges a published claim. Which principle of substantiation have they violated?
2. A brochure claims an implant delivers '40% fewer revisions,' but the underlying study measured a different population than the one targeted by the ad. Which substantiation rule is at risk?
3. Why is the standard 'clinicians tell us it works' insufficient for a substantiation file?
4. Select ALL correct answers about how promotional oversight is divided in the United States.
Select all the correct answers.
5. Select ALL correct answers describing the core purpose and structure of a substantiation file.
Select all the correct answers.
When the file fails: the cost
In the US, the FDA can issue an Untitled Letter or the heavier Warning Letter, demanding withdrawal of materials and, in serious cases, corrective communications to the audience that saw them. These are published in the Warning Letters database, so a promotional violation becomes permanently searchable. Competitors also file complaints, in the US often to the National Advertising Division (NAD) of BBB National Programs, whose rulings are fast, public and awkward to unwind.
Theranos is the outer limit: a company marketing hundreds of blood tests while publishing almost nothing peer-reviewed on the analyser behind them. When regulators finally inspected the Newark laboratory in 2016, its CLIA certificate was revoked and the company voided or corrected two years of test results. Elizabeth Holmes was convicted on four counts of fraud in January 2022 and sentenced to more than eleven years. No substantiation file existed to produce, which is why the claims survived so long internally: nobody had a document to disagree with.
The quieter version is more instructive. Babylon Health kept promoting triage performance while a clinician publicly logged safety concerns about the chatbot's outputs, which drew UK regulator attention in 2020. Once the original evidence has been challenged in a journal, every downstream asset resting on it is exposed, including the ones nobody remembers approving.
A simple worked check
You do not need statistics to sanity-test a comparative claim. Try this on any "X% better" line:
Claim: "40% fewer revisions"
Step 1 Absolute rates? Comparator 5.0% → Device 3.0%
Step 2 Absolute difference 5.0% - 3.0% = 2.0 percentage points
Step 3 Relative reduction 2.0 / 5.0 = 40%
Step 4 Ask: which framing does the brochure use, and is it clear?Both numbers are true. But "40% fewer" sounds far larger than "2 in 100 fewer". Reviewers increasingly expect the absolute figure alongside the relative one. If your file only supports the relative framing, flag it before someone else does.
Key takeaways
- One claim, one traceable source, assembled before publication and stored where a reviewer can open it unaided.
- Grade the reference. A poster, a subgroup analysis and "data on file" each support narrower wording than a pre-specified primary endpoint. Citing a review instead of the primary paper is how stale claims survive.
- Match evidence type to claim: clinical for patient outcomes, bench for engineering specs, head-to-head for any competitor mention. A bench result never implies a clinical benefit.
- Real evidence can still overshoot. Grail's test detects a signal; that is not the same sentence as reducing mortality, and the file has to know the difference.
- The pack has to survive translation, format compression and field decks. Most broken claims were compliant on the day they were approved.