A pilot that survives contact with fee earners
Six weeks. That is how long it took a well-funded legal research tool to go from "this will change how we work" in a partner demo to a login nobody used, at a mid-sized litigation firm that piloted it in late 2024. The demo answered a planted question beautifully. Week one of real matter work, an associate asked it a genuinely ambiguous question about conflicting appellate authority, got a confident but wrong synthesis, and told three colleagues by lunch. The tool was not reopened.
This is not a story about bad AI. It is a story about a bad pilot. The tool may have been fine. The way it was introduced guaranteed failure.
Why demos lie
Vendor demos are optimized for the demo, not for your matters. Sales engineers know which prompts work. They have tested the exact query on the exact document set beforehand. Fee earner (the billing-generating lawyer: partner, associate, or counsel) time is expensive and scarce, so firms often let a single polished session substitute for real testing.
The gap between demo and deployment shows up in three places:
- Data mismatch. The demo ran on clean, well-tagged documents. Your document management system (DMS, the repository that stores and version-controls matter files, e.g. iManage or NetDocuments) has fifteen years of inconsistent naming, duplicate drafts, and unstructured folders.
- Task mismatch. The demo answered a question chosen because the tool answers it well. Real fee earners ask the questions that are hardest, not easiest, because those are the ones they need help with.
- Trust mismatch. A partner watching a demo is a spectator. A partner relying on output to advise a client is exposed to professional liability if that output is wrong. The psychological stakes are entirely different.
Design the pilot around failure, not success
A pilot that survives contact with fee earners is built to surface failure fast, cheaply, and privately, before failure surfaces in front of a client or a judge.
Pick a bounded, real matter, not a demo sandbox
Choose one active but low-risk matter category, for example routine commercial lease review or first-pass due diligence review on a mid-size deal, and use the tool on live documents with a real deadline. Sandboxed "test documents" produce sandbox-quality feedback.
Assign a skeptic, not a champion
Every pilot has a natural enthusiast. Put a second, harder-to-convince fee earner on the same workflow. Adoption data collected only from believers is not adoption data.
Set a falsifiable success metric before you start
"Lawyers liked it" is not a metric. Better ones:
- Time from document receipt to first-draft issues list, measured against the same task done manually the prior quarter.
- Number of hallucinated citations per 100 queries, checked by a supervising associate (a hallucinationhallucinationA hallucination is when an AI model generates output that is fluent and confident but factually wrong, fabricated, or unsupported by its source data.View full definition → is a fabricated but plausible-sounding fact or citation the model generates with no basis in the source material).
- Percentage of outputs a partner accepts with zero edits versus needing substantive rework.
Pick two or three, measure them for four to six weeks, and write them down before the pilot starts so success is not redefined afterward to match whatever happened.
Match the tool category to the task, honestly
The module's core categories behave very differently under pressure:
| Category | What it's honestly good at | Where it breaks |
|---|---|---|
| Drafting | First-pass clauses, routine agreements from templates | Novel deal structures, jurisdiction-specific nuance |
| Contract review | Flagging deviations from a playbook at speed | Judgment calls on commercial risk tolerance |
| Legal research | Surfacing candidate authority fast | Confidently synthesizing conflicting or overruled case law |
| Discovery / eDiscovery | Sorting large document volumes by relevance | Edge-case privilege calls, nuanced responsiveness |
| Knowledge retrieval | Finding "have we done this before" internal precedent | Answering questions the firm's own documents don't cover |
A pilot fails fast when the task chosen plays to the category's actual strength. It fails slow and ugly when the task chosen plays to its weakness, because the tool looks fine until the one matter where it matters most.
Wire it into the DMS from day one, not at rollout
If the pilot runs on uploaded sample files instead of a live connection to the firm's DMS, you are testing a different, easier product. Real integration surfaces the messy stuff: version control conflicts, permission boundaries (should the tool see documents an associate is walled off from under an ethical screen?), and retrieval quality against actual folder chaos. The Sedona Conference publishes freely available frameworks on AI and eDiscovery that are useful background on how retrieval quality gets evaluated in practice.
Make the failure cheap to report
The lease-review tool at the firm above did not fail because it made one mistake. It failed because the associate had no easy, low-status way to flag the mistake, so she just stopped opening the tool. Build a two-click "this was wrong" feedback path, and make sure the pilot lead actually reads and triages it weekly. Silence from users is not satisfaction. It is usually abandonment in progress.
Knowledge check
1. Why does a successful vendor demo provide weak evidence that a legal AI tool will work well in real practice?
2. What is the core difference between the 'task mismatch' and 'data mismatch' problems described in the lesson?
3. According to the lesson, why is the 'trust mismatch' particularly significant compared to the other mismatches?
4. Select ALL correct answers about why the litigation firm's pilot failed despite a strong initial demo.
Select all the correct answers.
5. Select ALL correct answers describing sound principles for designing an AI pilot that can survive real fee earner use.
Select all the correct answers.
Change management: partners as the hard case
Associates adopt tools that save them hours. Partners adopt tools they trust with their name and their client relationship. Those are different psychological hurdles, and most pilots only solve the first one.
Three things move partner trust in practice:
- Show the citation trail, always. A research or review output with no traceable source is unusable to someone who has to defend the answer to a client or a court. Tools that show exactly which paragraph, clause, or case the output came from get trusted faster than tools that just assert an answer.
- Start with review, not generation. Partners are more comfortable with AI checking their associate's work than with AI producing work under their name unsupervised. Frame the pilot as a second set of eyes before framing it as a first draft.
- Let a respected partner, not IT, announce it. Adoption inside a partnership is a social process. A memo from the innovation committee gets skimmed. A senior litigation partner saying "I used this on the Henderson matter and it caught something I'd missed" gets acted on.
🎬 [VIDEO: "How Law Firms Are Actually Using AI in 2025" — youtube.com — a practitioner-level discussion of real deployment patterns and adoption friction inside law firms, useful context alongside this lesson]
A minimal pilot scorecard
A simple weekly tracking structure keeps the pilot honest:
Week | Queries run | Flagged errors | Partner sign-off rate | Time saved (est. hrs)
1 | 42 | 6 | 40% | 3
2 | 58 | 4 | 55% | 6
3 | 61 | 2 | 71% | 9If error rate is not falling and sign-off rate is not rising by week three or four, that is a signal to redesign the workflow or the task, not to extend the pilot hoping it improves on its own.
Key Takeaways
- Demos are optimized to succeed; pilots must be designed to surface failure early, on real matters, with real deadlines and messy documents pulled through the actual DMS.
- Match the tool category (drafting, review, research, discovery, knowledge retrieval) honestly to its known strengths; task choice determines whether the pilot tests the tool's best case or its worst case.
- Set falsifiable, quantitative success metrics before the pilot starts, include a deliberate skeptic user alongside an enthusiast, and build an easy feedback loop so errors get reported instead of quietly avoided.
- Partner trust is a distinct problem from associate time savings: citation traceability, a "review not generate" framing, and peer advocacy from a respected partner move adoption more than any feature list.
- A tool that dies quietly in week two almost always died from pilot design, not from model quality; redesigning the pilot is usually cheaper than switching vendors.