# Evaluating and comparing hospital AI vendors
Two hospitals buy the same sepsis-prediction algorithm. One cuts sepsis mortality; the other quietly switches it off after six months because nurses stopped trusting the alerts. Same software, opposite outcomes. The difference was almost never the model itself. It was integration, validation fit, and how the tool changed clinical workflow.
That gap is why you need a structured scorecard, not a vendor demo. This lesson walks through comparing three sepsis-prediction vendors head-to-head using four evaluation lenses that actually predict success.
Sepsis is a life-threatening response to infection. Every hour of delayed treatment raises mortality, so early prediction is a high-value AI use case. It is also the most instructive one to evaluate, because the field already had a public failure.
The Epic Sepsis Model (a prediction tool embedded in the widely used Epic electronic health record, or EHR: the digital system storing patient charts) was independently studied in a 2021 *JAMA Internal Medicine* paper. Researchers at Michigan found it performed far worse in real use than the vendor's own numbers suggested, missing many sepsis cases while flooding clinicians with alerts. See the study summary here: JAMA Internal Medicine, 2021.
The lesson: vendor-reported accuracy is a starting point, not evidence.
We will score three hypothetical but realistic vendor profiles on a 1 to 5 scale across four lenses. The vendors:
An AI model is useless if it cannot see live patient data and cannot surface predictions where clinicians already work. Integration is usually the single biggest driver of success or failure.
Key questions:
| Vendor | Integration profile | Score |
|--------|--------------------|-------|
| A (native) | Runs inside EHR, real-time, alerts in existing workflow | 5 |
| B (specialist) | FHIR integration, mature but requires interface build | 4 |
| C (startup) | APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.View full definition → only, alerts in separate dashboard, batch data | 2 |
Here you separate marketing from medicine. Ask for evidence in this hierarchy, strongest first:
1. Prospective study: tested on live patients going forward, ideally at multiple sites.
2. External validation: tested at hospitals other than where it was built.
3. Retrospective internal: tested only on the vendor's historical data (weakest, and what most vendors show first).
Two metrics matter, and vendors love to hide behind one:
A model can have high sensitivity and terrible PPV, meaning it catches most cases but drowns staff in false alarms. That is exactly what caused the Epic model backlash and what drives alert fatigue (clinicians tuning out warnings because too many are wrong).
Also ask: was the model FDA-cleared as Software as a Medical Device (SaMD)? In the US, the Food and Drug Administration regulates many diagnostic AI tools. In Europe, the equivalent is CE marking under the Medical Device Regulation (MDR), plus obligations arriving under the EU AI Act, which classifies most clinical decision AI as high-risk. Clearance is not proof of real-world performance, but its absence is a red flag for a diagnostic claim.
| Vendor | Evidence | Score |
|--------|----------|-------|
| A (native) | Retrospective internal only, no external validation | 2 |
| B (specialist) | Prospective multi-site study, FDA-cleared | 5 |
| C (startup) | Retrospective, impressive AUC, no clinical trial | 2 |
Note: AUC (area under the curve) is a common accuracy summary from 0.5 to 1.0. A high AUC on retrospective data does not guarantee bedside value.
How the model lives and updates matters for safety and IT burden.
| Vendor | Deployment | Score |
|--------|-----------|-------|
| A (native) | Cloud, tied to EHR update cycle, no separate drift dashboard | 3 |
| B (specialist) | Hybrid, quarterly recalibration, drift monitoring included | 5 |
| C (startup) | Cloud only, unclear drift process | 2 |
Knowledge check
1. Two hospitals deployed the same sepsis-prediction algorithm, yet one reduced mortality while the other abandoned it. What does this scenario primarily illustrate about evaluating AI vendors?
2. Why does the lesson argue that a structured scorecard is preferable to relying on a vendor demo when comparing AI tools?
3. The Epic Sepsis Model performed far worse in independent study than the vendor's own numbers suggested. What general principle should a buyer draw from this?
4. Select ALL correct answers. Why is EHR integration a critical evaluation lens for a hospital AI tool?
Select all the correct answers.
5. Select ALL correct answers. Why is sepsis prediction described as an especially instructive test case for evaluating AI vendors?
Select all the correct answers.
Not the license fee. The full cost of running AI over three years. Stay focused on AI-specific costs, not general hospital finance.
TCO components:
That last item is routinely underestimated. Alert tuning and clinician training often cost as much as the software in year one.
#### Worked TCO example (illustrative estimates only)
These are made-up round numbers to show the method, not real vendor prices. Assume a 300-bed hospital, three-year horizon, Vendor B:
Three-year TCO = 450,000 + 120,000 + 120,000 + 150,000 = $840,000
Now compare against benefit. Suppose the tool contributes to earlier treatment for an estimated 20 additional sepsis patients per year, and each averted deterioration saves an estimated $15,000 in avoided ICU days (again, illustrative):
Net over three years: 900,000 minus 840,000 = $60,000 positive, with most value coming from clinical outcomes, not headcount savings. The point is not the number. It is that a cheaper license (Vendor C) can lose to a pricier one (Vendor B) once integration, drift monitoring, and change management are included.
Sum the four lenses (max 20):
| Vendor | Integration | Validation | Deployment | TCO fit | Total |
|--------|:-:|:-:|:-:|:-:|:-:|
| A (native) | 5 | 2 | 3 | 4 | 14 |
| B (specialist) | 4 | 5 | 5 | 3 | 17 |
| C (startup) | 2 | 2 | 2 | 5 | 11 |
Weight the lenses to your context. A hospital with a stretched IT team may weight integration highest, making Vendor A more attractive despite weaker evidence. A large academic center may refuse anything without prospective validation, ruling out A and C regardless of price.
The scorecard does not pick the winner for you. It forces the tradeoffs into the open and stops a slick demo from deciding a clinical safety purchase.
🎬 [VIDEO: "How to Evaluate Clinical AI Tools" - youtube.com - a plain-language walkthrough of validation and integration pitfalls in hospital AI procurement]
Even the best-scoring vendor will not deliver value in isolation. Three sobering realities: