+150 XP

Building a defensible pilot for AI in a practice group

A managing partner at a 200-lawyer firm once approved a firm-wide AI rollout because one senior associate said a contract-review tool "felt faster." Eighteen months later, adoption had stalled below 20%, nobody could say whether error rates had changed, and the tool's renewal came up for a vote with zero usable evidence either way. That is what happens when a pilot is designed to produce a good story instead of a good dataset.

This lesson is about building the alternative: a 90-day pilot inside one practice group (real estate or litigation, for concreteness) that generates evidence a skeptical partner committee will actually trust.

Why most AI pilots fail as evidence

Law firm pilots typically fail not because the tool is bad, but because the pilot design cannot answer three questions:

  1. Compared to what baseline?
  2. Measured how, by whom, on what sample?
  3. Would the result replicate with a different associate, matter, or month?

An anecdote ("the AI drafted a solid indemnification clause") answers none of these. A defensible pilot is built backward from the questions the partnership will ask before committing budget.

Step 1: Pick a bounded, measurable workflow

Do not pilot "AI in litigation." Pilot one task with a clear output and a known cost today.

Litigation example: first-pass privilege review on a document set of 5,000-20,000 documents, using an AI classification tool layered on top of existing e-discovery platforms (Relativity, Everlaw). Privilege review has a defined unit cost (associate or contract attorney hours per 1,000 documents) and a defined error consequence (a privilege waiver, which is a real, sanctionable risk under attorney-client privilege doctrine).

Real estate example: AI-assisted lease abstraction, extracting rent escalation clauses, renewal options, and assignment restrictions from a portfolio of commercial leases ahead of a due diligence deadline. Abstraction has a known per-lease cost baseline from paralegal or junior associate time.

Bounding the workflow lets you measure against a real baseline instead of a vague "productivity gain."

Step 2: Build the baseline before touching the tool

This is the step most pilots skip, and it is the one that makes the results defensible.

For 2 to 3 weeks before the AI tool is introduced, track the current process on a comparable sample:

  • Time per document/lease (hours, logged, not estimated after the fact)
  • Error rate on a partner-reviewed sample (e.g., 10% of output re-checked by a senior lawyer)
  • Cost per unit, using realized billing rates or internal cost rates, not rack rates

Without this baseline, any "40% faster" claim from the AI phase is unfalsifiable. With it, you have a control group of one, which is thin but far better than nothing.

Step 3: Run the pilot with a fixed evaluation protocol

Over the 90 days, structure it in three four-week blocks:

Weeks 1-4: Supervised parallel run. Every AI output is checked by a human before it counts. This measures accuracy, not speed, because lawyers are still fully reviewing.

Weeks 5-8: Spot-check run. Reduce human review to a fixed sample (e.g., 20% of AI outputs), simulating real deployment. This is where you measure realistic speed gains.

Weeks 9-12: Steady state with incident logging. Full deployment on the pilot workflow, with a mandatory log of every "AI got this materially wrong" event, however minor.

Track four metrics consistently across all three blocks:

MetricWhat it tells the partnership
Time per unit (hours)Real throughput gain, not perceived speed
Error/exception rateWhether accuracy holds outside supervised conditions
Cost per unit (fully loaded)Whether savings survive licensing and review overhead
Associate/partner override rateTrust signal: how often humans reject AI output

The override rate matters more than most firms realize. A tool with a 90% acceptance rate in week 2 and a 60% acceptance rate in week 10 is a red flag, not a maturing tool; it usually means users stopped checking carefully, not that the tool got worse.

Step 4: Price it like a lawyer, not a vendor

Vendors quote per-seat licensing (commonly in the low hundreds of dollars per user per month for enterprise legal AI tools, as an industry estimate for 2025-2026, varying widely by vendor and contract size). That number alone tells the partnership nothing about ROI.

Worked example (illustrative, not a real vendor's pricing):

  • Baseline: paralegal lease abstraction, 45 minutes/lease, loaded cost estimate $60/hour → $45/lease
  • Pilot (steady state, weeks 9-12): 20 minutes/lease review time + tool cost allocated per lease
  • Tool cost: $400/seat/month, 3 seats, 200 leases/month → $6/lease allocated
  • New cost: (20 min × $60/hr) + $6 = $20 + $6 = $26/lease
  • Saving: $19/lease, or about 42%, on this workflow only

This is a hypothetical illustration to show the calculation method, not a benchmark to expect. Every firm must run its own numbers with its own rates and its own tool pricing.

The critical discipline: include review time, not just AI processing time, and include the license cost allocated to actual volume, not list price. Pilots that only measure "how fast did the AI run" wildly overstate ROI.

Step 5: Build the risk case alongside the ROI case

Partners on a pilot review committee will ask about downside before upside. Prepare answers grounded in real regulatory context:

  • Confidentiality and privilege: Under ABA Model Rule 1.6 (duty of confidentiality) and its state equivalents, uploading client documents to a third-party AI tool raises questions about data handling, vendor contracts, and whether the tool trains on client data. Confirm the vendor's data policy in writing.
  • Competence duty: ABA Model Rule 1.1 (and its 2012 comment on technology competence) means lawyers must understand a tool well enough to supervise it, not delegate judgment to it.
  • EU context: firms with European matters should track the EU AI Act, which classifies certain AI uses by risk tier; most current legal-drafting and review tools fall outside the "high-risk" category but firms should document their own classification reasoning.
  • Malpractice exposure: a documented human-review step in the pilot protocol is itself a risk-mitigation artifact, useful if a client ever challenges the process.

Documenting these three points *inside the pilot report* turns "trust me" into "here is our reasoning."

Vérification des acquis

1. According to the lesson, why did the firm-wide AI rollout described in the opening example ultimately fail to provide useful evidence for the renewal decision?

2. What is the main reason the lesson advises against piloting something as broad as 'AI in litigation'?

3. Why does privilege review serve as a good example of a 'bounded, measurable workflow' for a litigation AI pilot?

CHOIX MULTIPLES

4. Select ALL correct answers: which questions must a defensible AI pilot be able to answer, according to the lesson?

Sélectionnez toutes les réponses correctes.

CHOIX MULTIPLES

5. Select ALL correct answers: which of the following are characteristics of a well-bounded pilot workflow, as illustrated by the litigation and real estate examples?

Sélectionnez toutes les réponses correctes.

Step 6: Present results as a decision memo, not a demo

The final deliverable should look like a one-page table plus a one-page narrative, not a slide deck of screenshots. Include:

  • Baseline vs. pilot metrics (time, cost, error rate, override rate) by block
  • Total pilot cost (licenses, training time, review overhead)
  • Net cost/hour saved, with the calculation shown
  • Named risks and mitigations, tied to specific ethics rules
  • A recommendation: expand, extend the pilot, or stop, with the threshold that would change the recommendation

This format survives partner committee scrutiny because every number traces back to a method, not a feeling.

🎬 [VIDEO: "How Law Firms Are Actually Using AI in 2025" — youtube.com — search for recent legal-industry panel discussions (e.g., from Legalweek or ABA TECHSHOW sessions) covering real deployment data across litigation and transactional practices]

Common pitfalls to design out

  • Selection bias: letting the most enthusiastic associate self-select into the pilot. Rotate 3 to 5 users of varying seniority instead.
  • No stop condition: define in advance what error rate or override rate would trigger pausing the pilot.
  • Comparing to an unrealistic baseline: firms sometimes benchmark against a slow, under-resourced status quo, inflating apparent AI gains. Use the last 90 days of actual practice, not an idealized one.
  • Ignoring associate development cost: junior lawyers learn drafting and issue-spotting partly through the manual work AI now automates. A defensible pilot notes this training tradeoff explicitly, it is a real, if hard-to-quantify, cost.

For a useful framework on evaluating AI tools generally (applicable beyond legal), see NIST's AI Risk Management Framework, which offers a structured way to think about measuring and documenting AI performance and risk.

Key Takeaways

  • A defensible pilot starts with a bounded, measurable workflow (privilege review, lease abstraction) and a pre-AI baseline measured on the same team, same time period.
  • Structure the 90 days in three blocks (supervised, spot-checked, steady state) and track time, error rate, cost per unit, and override rate consistently across all three.
  • Calculate ROI with fully loaded costs, review time included, license cost allocated per unit, not vendor list price, never processing speed alone.
  • Ground the risk discussion in named rules (ABA Model Rules 1.1 and 1.6, EU AI Act where relevant), not general caution, so partners can evaluate real exposure.
  • Deliver a decision memo with traceable numbers and a stated stop/expand threshold, not a demo of impressive outputs.