Reading a legal AI vendor's evaluation like a discovery deposition
A vendor's slide says "94% accuracy on contract review." You don't ask what it means. You ask: 94% of what, measured how, on whose documents, compared to what baseline. That's the entire lesson.
Law firms and legal departments are buying AI contract review, e-discovery, and legal research tools at a fast clip in 2026. Most buyers accept the headline number and move to pricing. That's the equivalent of letting an expert witness's CV substitute for cross-examination. This lesson gives you the questions to ask instead.
Where legal AI actually earns its keep
Before evaluating any claim, know where AI plausibly adds value in a law firm's workflow, because that shapes what a "good" benchmark looks like.
Contract review and due diligence. Tools like Harvey, Ironclad, Kira Systems, and Luminance extract clauses (indemnification, termination, change of control) from contracts at speeds no associate can match. This is the most mature use case, because the task is narrow: find and classify text, not exercise judgment.
E-discovery. Technology-assisted review (TAR), a method using machine learning to rank documents by relevance in litigation, has been court-recognized since *Da Silva Moore v. Publicis Groupe* (2012, US federal case). It remains the clearest case of AI outperforming manual review on cost and consistency.
Legal research. Tools like Casetext (acquired by Thomson Reuters in 2023) and Lexis+ AI generate case summaries and draft memos using large language models (LLMs), AI systems trained on text to generate human-like responses. This is higher-risk: LLMs can "hallucinate," meaning they generate plausible but fabricated citations. This caused real sanctions, notably *Mata v. Avianca* (2023, SDNY), where lawyers cited fake cases produced by ChatGPT.
Drafting assistance. Generating first-draft NDAs, memos, or discovery requests. High adoption, but value depends entirely on how much partner/associate review time it actually saves versus how much it shifts to fact-checking.
The pattern: AI performs best on narrow, well-defined extraction and classification tasks, and remains riskiest on generative tasks requiring legal judgment and citation accuracy.
The deposition questions: interrogating the 94%
Treat every vendor accuracy claim like a witness statement. Here's your cross-examination checklist.
1. Accuracy of what, exactly?
"94% accuracy" on contract review could mean: percentage of clauses correctly identified, percentage of contracts with zero errors, or F1 score (a statistical measure combining precision and recall, explained below). These are wildly different claims. Ask for the precise metric definition in writing.
2. Precision vs. recall: the confusion that hides bad tools
- Precision: of the clauses the AI flagged as "problematic," what percentage actually were? Low precision means associates waste time chasing false alarms.
- Recall: of all the actually problematic clauses in the document set, what percentage did the AI catch? Low recall means real risk gets missed, silently.
A vendor can hit 94% precision while missing half the risky clauses (poor recall), and the headline number will still look great.
F1 score is the harmonic mean of precision and recall, useful because it punishes tools that are lopsided in either direction:
F1 = 2 × (Precision × Recall) / (Precision + Recall)
Example:
Precision = 0.94 (94% of flagged clauses are real issues)
Recall = 0.60 (only 60% of actual issues get flagged)
F1 = 2 × (0.94 × 0.60) / (0.94 + 0.60) = 1.128 / 1.54 ≈ 0.73
So a "94% accurate" tool can carry a 73% F1 score once recall is factored in.Always ask for both numbers, not just the flattering one.
3. Tested on documents like yours?
A model benchmarked on standardized US commercial leases will underperform on bespoke cross-border joint venture agreements or EU-law-governed supply contracts. Ask: what document types, what jurisdictions, what governing law, how much contract-to-contract variation existed in the test set. A benchmark on homogenous templates does not predict performance on your messy, negotiated, multi-jurisdiction contract portfolio.
4. Compared to what baseline?
94% accuracy sounds impressive until you learn that experienced paralegals hit 97% on the same task, or that a simple keyword search hits 85%. Demand the human benchmark used for comparison, and how many reviewers, and their seniority.
5. Who validated the test set (the "ground truth")?
Ground truth means the correct answers the AI's outputs are compared against. If the vendor's own team labeled the ground truth, that's a conflict of interest, the equivalent of an expert witness grading their own exam. Ask whether the labels were produced by independent lawyers, and whether inter-annotator agreement (how often independent human reviewers agreed with each other) was measured. If two experienced lawyers disagree on the "correct" answer 20% of the time, no AI benchmark above that ceiling is fully meaningful.
6. Statistical significance and sample size
"94% accuracy" on a test set of 40 documents tells you almost nothing. Ask for sample size and confidence intervals. A useful public reference for how legal AI benchmarking should be approached rigorously is Stanford's RegLab and CodeX legal AI research, which has published independent analyses of legal AI tool performance, including studies showing legal AI hallucinationAI hallucinationUne hallucination, c'est lorsqu'un modèle d'IA produit une réponse fluide et assurée mais factuellement fausse, inventée, ou non étayée par ses données sources.Voir la définition complète → rates vary significantly by vendor and task.
Vérification des acquis
1. A vendor claims '94% accuracy on contract review.' What is the most important follow-up before evaluating this claim?
2. Why is contract review described as the most mature use case for legal AI, while legal research (LLM-generated case summaries) is described as higher-risk?
3. A law firm is deciding whether to trust an AI tool's output without independent verification. Based on the lesson's framework, which task would warrant the LEAST additional scrutiny before relying on the output?
4. Select ALL correct answers describing why technology-assisted review (TAR) in e-discovery is considered a well-established, lower-risk AI application compared to LLM-based legal research.
Sélectionnez toutes les réponses correctes.
5. Select ALL correct answers about questions a buyer should ask when evaluating a legal AI vendor's benchmark claim, based on the discovery-deposition framing in the lesson.
Sélectionnez toutes les réponses correctes.
Adoption and ROIROIReturn on Investment : le rapport entre le profit net et le coût d'un investissement. Un ROI de 300 % signifie que chaque dollar investi en rapporte 3.Voir la définition complète →: realistic expectations
Even a well-validated tool doesn't automatically produce ROIROIReturn on Investment : le rapport entre le profit net et le coût d'un investissement. Un ROI de 300 % signifie que chaque dollar investi en rapporte 3.Voir la définition complète → (return on investmentreturn on investmentReturn on Investment : le rapport entre le profit net et le coût d'un investissement. Un ROI de 300 % signifie que chaque dollar investi en rapporte 3.Voir la définition complète →). Three factors determine whether it does.
Billable hour tension. Traditional law firm economics reward hours billed. A tool that cuts contract review from 8 hours to 2 hours can reduce revenue per matter under hourly billing, unless the firm shifts to value-based or fixed fees for that work. This is a real, documented friction in legal AI adoption, not a hypothetical. Firms that see the clearest ROIROIReturn on Investment : le rapport entre le profit net et le coût d'un investissement. Un ROI de 300 % signifie que chaque dollar investi en rapporte 3.Voir la définition complète → have often already moved parts of their practice to alternative fee arrangements.
Review burden doesn't disappear, it shifts. Someone still has to check the AI's output, especially for citation accuracy in research tools. A partner spending 30 minutes verifying an AI draft that used to take an associate 3 hours to write is real time savings, but only if the verification time is honestly counted. Vendors' ROIROIReturn on Investment : le rapport entre le profit net et le coût d'un investissement. Un ROI de 300 % signifie que chaque dollar investi en rapporte 3.Voir la définition complète → case studies rarely show the verification cost.
Realistic estimate ranges (2025-2026, industry-reported, treat as estimates): Contract review time reductions of 40-60% are commonly cited by vendors like Harvey and Ironclad in customer case studies; independent, peer-reviewed confirmation at that scale is limited. E-discovery cost reductions from TAR are better documented, with legal industry sources citing reductions of 20-60% in review costs versus fully manual linear review, varying heavily by matter size and document volume. Treat single-number ROIROIReturn on Investment : le rapport entre le profit net et le coût d'un investissement. Un ROI de 300 % signifie que chaque dollar investi en rapporte 3.Voir la définition complète → claims from any one vendor's case study with the same skepticism as their accuracy claims: ask about the specific matter type, document volume, and who measured the "before" baseline.
Governance matters as much as accuracy. In both the US and EU, courts have begun requiring disclosure when AI tools are used in filings, and the EU's AI Act (Regulation (EU) 2024/1689, entered into force August 2024) classifies certain AI use in the administration of justice as "high-risk," triggering documentation, human oversight, and risk management obligations for providers and, in some cases, deploying law firms. Ask any vendor how their tool is classified and what compliance documentation they provide.
🎬 [VIDEO: "AI in Legal Practice: Promise and Peril" — youtube.com — search for Stanford Law School or American Bar Association panel discussions on generative AI risks in legal practice, covering hallucinationhallucinationUne hallucination, c'est lorsqu'un modèle d'IA produit une réponse fluide et assurée mais factuellement fausse, inventée, ou non étayée par ses données sources.Voir la définition complète → cases and courtroom sanctions]
Key Takeaways
- Never accept a single accuracy number. Demand precision, recall, and F1 score separately, and ask how "accuracy" was defined.
- Insist the benchmark used documents matching your jurisdiction, contract type, and complexity, not clean template data.
- Check who created the ground truth labels (independent lawyers, not the vendor) and what the sample size was.
- ROIROIReturn on Investment : le rapport entre le profit net et le coût d'un investissement. Un ROI de 300 % signifie que chaque dollar investi en rapporte 3.Voir la définition complète → depends on billing model and honest accounting of human verification time, not just the vendor's time-saved percentage.
- Know the regulatory backdrop: EU AI Act high-risk classification and US court disclosure rules on AI use are becoming standard evaluation criteria, not optional extras.