# Evaluating AI vendor claims and demos
"Watch this. One click, and we've turned a 40 minute city council meeting into a published article, tagged, SEOSEOSearch Engine Optimization: the practice of improving your pages' natural (unpaid) rankings in search engine results pages to attract more organic traffic.Voir la définition complète → optimized, and ready for review, in under 90 seconds."
The room nods. The video shows a clean transcript, a punchy headline, a summary with quotes pulled correctly. A local news group signs a pilot contract that week.
Three months later, the tool mangles a councilmember's name in every story, invents a budget figure that was never said, and needs a human editor to rewrite half the output anyway. The pilot quietly dies.
This happens constantly across media AI procurement. Not because vendors lie outright, but because demos are optimized to persuade, not to inform. This lesson teaches you to interrogate a demo the way you'd interrogate a financial audit: line by line, assumption by assumption.
Vendors don't need to fabricate results. They just need to show you the best case.
Cherry-picked inputs. The council meeting in the demo was probably chosen because it had clear audio, a single speaker at a time, and simple subject matter. Your actual meetings have crosstalk, regional accents, and technical jargon about zoning ordinances.
Benchmark mismatch. A vendor claiming "95% transcription accuracy" is likely citing a public benchmark like LibriSpeech (clean, read-aloud English audiobooks) or a curated internal test set. That number tells you almost nothing about performance on a noisy press conference or a heavily accented interview.
Survivorship in the reel. You're seeing the outputs that made the cut, not the failed attempts edited out before the demo was recorded.
None of this is fraud. It's marketing. Your job is to ask the questions that separate the marketing layer from the underlying capability.
Let's go back to that opening line and pull it apart.
> "One click, and we've turned a 40 minute meeting into a published article."
"One click": What happened before the click? Was the audio pre-cleaned? Was there a custom prompt tuned specifically for council meeting formats? Ask: *what setup, templates, or fine-tuningfine-tuningFine-tuning adapts a pre-trained model to a specific task or domain by continuing training on a smaller, targeted dataset, improving accuracy and style for that use case.Voir la définition complète → does this require before it works on our content?*
"40 minute meeting": Which meeting? Ask to run the tool live on a meeting *you* choose, ideally one with known difficulty (overlapping speakers, poor audio, local names).
"Published article": Published by whom, under what editorial standard? Ask: *what percentage of output required human correction before it met our publication standard, measured across a representative sample, not one clip?*
"90 seconds": Processing speed is usually the least important variable. Ask: *what's the cost and time of the human review step that follows?* That's the real cycle time.
1. "Can I bring my own data right now?" A vendor confident in their product will let you test on your actual archive, your actual audio quality, your actual house style, live in the room.
2. "What's the benchmark, exactly, and how does it relate to our content?" Ask for the dataset name, the metric definition, and the date measured. Vague answers ("industry-leading accuracy") mean the number wasn't chosen to inform you.
3. "What data do you need from us to reach these numbers?" Many tools need weeks of your archive to fine-tune (adapt a general model to your specific content) before performance resembles the demo. That's a hidden cost and a hidden delay.
4. "What does failure look like, and how often does it happen?" Ask for error rates, not just success rates. A tool that's 90% accurate is wrong 1 time in 10, ask what "wrong" looks like in practice: a misspelled name, or a fabricated quote.
5. "Who else in our sector uses this in production, not pilot, and can I talk to them?" Reference customers in similar workflows (not just similar industry) reveal operational reality.
Say two vendors both claim "around 95% accuracy" (word error rate, WER, the standard metric for transcription quality: the percentage of words transcribed incorrectly, subtracted from 100).
WER is calculated as:
WER = (Substitutions + Insertions + Deletions) / Total words in reference transcriptVendor A tested on studio-quality podcast audio. Vendor B tested on a mix including phone interviews and outdoor press conferences.
If your newsroom's actual audio is 60% studio quality and 40% field recordings, neither number tells you your real cost. Run your own test:
If Vendor A shows 95% on studio audio but drops to 80% on field audio, and 40% of your work is field audio, your blended real-world accuracy is roughly:
(0.95 × 0.6) + (0.80 × 0.4) = 0.57 + 0.32 = 0.89, or about 89%, not 95%.
That 6 point gap is the difference between "light editing" and "re-transcribing half the interview." It changes your staffing math and your ROIROIReturn on Investment: the ratio of net profit to the cost of an investment. A 300% ROI means each dollar invested returns $3.Voir la définition complète → (return on investmentreturn on investmentReturn on Investment: the ratio of net profit to the cost of an investment. A 300% ROI means each dollar invested returns $3.Voir la définition complète →) case entirely.
Vérification des acquis
1. A vendor cites a 95% transcription accuracy figure from a benchmark like LibriSpeech. Why might this number be a poor predictor of performance on your organization's actual audio?
2. What is the most accurate way to characterize why AI vendor demos often mislead buyers?
3. A news director wants to evaluate a vendor's claim more rigorously before signing a contract. Which approach best reflects the lesson's core recommendation?
4. Select ALL correct answers describing reasons a polished AI demo can fail to predict real-world performance.
Sélectionnez toutes les réponses correctes.
5. Select ALL correct answers about the relationship between vendor honesty and misleading demos.
Sélectionnez toutes les réponses correctes.
Data preparation. Fine-tuningFine-tuningFine-tuning adapts a pre-trained model to a specific task or domain by continuing training on a smaller, targeted dataset, improving accuracy and style for that use case.Voir la définition complète → a model on your style guide or archive takes engineering time and clean historical data. If your archive is inconsistently tagged, that's weeks of prep, not a checkbox.
Integration. A generative AI (AI that produces new text, image, or audio content rather than just classifying existing content) tool that "plugs into your CMS" (content management system) still needs custom APIAPIApplication Programming Interface: a standardised interface that lets applications communicate and exchange data without knowing each other's internal workings.Voir la définition complète → work in most newsroom stacks, which vary widely.
Review workflow redesign. If AI drafts need mandatory human sign-off, as most responsible newsroom policies require, you need new editorial roles and time budgets, not just a new tool.
Ongoing drift. Model performance can degrade as your content changes (new beat, new format, new speaker patterns) or as the vendor updates their underlying model without notice. Ask about change logs and your right to re-test after updates.
The Associated Press's guidelines on generative AI are a useful public reference for what responsible review workflows look like in practice, useful benchmark language when negotiating vendor contracts.
Before any pilot, score vendors on:
| Criterion | What to check |
|---|---|
| Test transparency | Will they test live on your data? |
| Benchmark relevance | Do published numbers match your content type? |
| Failure disclosure | Do they share error rates and failure modes, not just averages? |
| Total cost | Does the quote include integration, fine-tuningfine-tuningFine-tuning adapts a pre-trained model to a specific task or domain by continuing training on a smaller, targeted dataset, improving accuracy and style for that use case.Voir la définition complète →, and review labor? |
| Reference quality | Can you speak to a production user in a comparable workflow? |
A vendor who scores well on transparency, even with modest headline numbers, is usually a safer bet than one with spectacular numbers and evasive answers.
🎬 [VIDEO: "How AI Transcription Actually Works (and Where It Fails)" - youtube.com - search for recent explainer content on word error rate and speech recognition limitations from a reputable AI or journalism education channel, useful for a non-technical primer before a vendor meeting]