+150 XP

Who owns the mistake when AI gets it wrong

A mid-size accounting firm in early 2024 used an AI research tool to draft a client memo on a tax position. The tool cited case law that did not exist. A junior associate, trusting the output, sent it to a partner who signed off without checking the citations. The client relied on it. This is not a hypothetical: courts in the US have sanctioned lawyers multiple times since 2023 for exactly this pattern, most notably in *Mata v. Avianca* (Southern District of New York, 2023), where attorneys were fined for submitting ChatGPT-generated fake citations. Professional services firms now face a version of this risk on every engagement touched by AI: valuations, audit workpapers, tax opinions, actuarial models, due diligence reports.

The question this lesson answers: when AI contributes to a professional judgment that turns out wrong, who is liable, and what structure keeps the firm defensible.

Why this is different from ordinary tool error

Professional liability (malpractice exposure tied to the standard of care owed to a client) has always assumed a human expert exercised judgment. AI changes three things:

  • Opacity: many AI outputs (especially from large language models, LLMs) don't show their reasoning, so reviewers can't easily spot the flaw the way they'd spot a junior's arithmetic error.
  • Confidence without competence: AI tools produce fluent, authoritative-sounding text regardless of accuracy. This is often called "hallucination," the generation of plausible but false content.
  • Scale: one bad prompt template or one flawed model can propagate the same error across hundreds of client files before anyone notices.

Regulators and courts have not created a separate "AI defense" or "AI liability shield." The existing standard of care still applies: a professional is expected to exercise the same competence whether they used AI or not. The American Bar Association's Formal Opinion 512 (2024) on generative AI made this explicit for lawyers: competence, confidentiality, and supervision duties apply fully to AI-assisted work. Expect equivalent guidance from accounting bodies (AICPA) and actuarial and valuation associations to keep tightening through 2026.

Mapping the liability chain

When something goes wrong, four parties can be in the liability chain, and firms should map this before deployment, not after an incident.

  1. The firm and the signing professional. Malpractice claims target the person or firm who delivered the opinion, filing, or valuation, regardless of what tool produced the draft. Signing off means owning it.
  2. The AI vendor. Contracts with vendors like Microsoft (Copilot), Thomson Reuters (CoCounsel), or Workiva typically disclaim liability for output accuracy through terms of service. Vendors sell "assistive" tools, not opinions.
  3. The client, if they misused the tool themselves or ignored disclosed limitations.
  4. Insurers, whose professional indemnity (errors and omissions, E&O) policies increasingly ask explicit questions about AI use during underwriting.

The practical result: liability defaults back to the firm almost every time. Vendor disclaimers are broad and largely untested but generally enforceable for tool-level errors. This means the firm's internal controls, not the AI vendor's promises, are the real risk mitigant.

The human sign-off: what actually counts as one

"A human reviewed it" is not a defense on its own. Regulators and courts are starting to ask for evidence of *meaningful* review. Build sign-off around three tests:

  • Traceability: can the reviewer see what the AI generated versus what a human added or changed?
  • Independent verification: were citations, figures, or calculations checked against a primary source, not just checked for plausibility?
  • Competence to catch the error: was the reviewer senior enough, and given enough time, to actually catch the class of error the tool is known to make?

A partner glancing at a 40-page AI-drafted valuation report for five minutes before signing does not meet this bar. A structured checklist tied to specific known failure modes (numerical fabrication, outdated regulatory citations, mismatched comparables in a valuation) does.

The documentation trail: what to keep

If a claim arises two years later, the firm needs to reconstruct exactly what happened. Minimum documentation:

  • Which AI tool and model version was used (model versions change silently; GPT-4 outputs in 2024 differ from GPT-4o or later versions).
  • What prompt or input was given.
  • The raw AI output, unedited.
  • What the human reviewer changed and why.
  • Who signed off and on what date.
engagement_id: 2026-0417
tool: CoCounsel_v3
task: precedent_summary
raw_output_hash: 8f2a1c...
human_edits: 3 citations removed (unverifiable), 1 paragraph rewritten
reviewer: J. Alvarez, Senior Associate
signoff: M. Chen, Partner, 2026-03-11
verification_method: manual check against Westlaw

This kind of audit log turns "we used AI" from a liability admission into evidence of a controlled process. The EU's AI Act (Regulation 2024/1689), which began phasing in through 2025 and 2026, pushes in the same direction for "high-risk" AI systems: it requires documentation, logging, and human oversight capable of intervening, particularly relevant for AI used in credit, employment, or legal-adjacent decisions that some professional services touch.

The insurance conversation

Professional indemnity insurers have moved fast. By 2025, major E&O underwriters (including specialty insurers like Beazley and Chubb) were adding explicit AI-use questionnaires to renewal applications, asking:

  • Do you use generative AI in client deliverables?
  • Do you have a policy governing AI use and human review?
  • Have you had any AI-related incidents or near-misses?

Firms that answer vaguely or admit to no policy risk higher premiums or exclusions. A firm with documented guardrails (this lesson's checklist, essentially) is a materially better underwriting risk. Some insurers, per industry commentary from brokers like Marsh, are starting to treat "no AI governance policy" the way they once treated "no cybersecurity policy": a red flag that triggers higher scrutiny or surcharge.

Practical move: before renewal, get the firm's actual AI use inventoried and put a one-page governance policy in front of the broker. It changes the negotiating position.

Knowledge check

1. In the Mata v. Avianca pattern, what was the core professional failure that led to sanctions?

2. Why does AI hallucination pose a harder review problem than a typical junior staff arithmetic error?

3. According to the standard of care discussed in this lesson, how does using AI change a professional's liability exposure?

MULTIPLE CHOICE

4. Select ALL correct answers describing why AI errors are structurally different from ordinary tool errors in professional services.

Select all the correct answers.

MULTIPLE CHOICE

5. Select ALL correct answers about what happened in the accounting firm scenario described in the lesson.

Select all the correct answers.

Guardrails to run before deployment

A pre-deployment checklist, adaptable across audit, tax, legal, actuarial, and valuation work:

  1. Classify the use case by risk. Drafting an internal summary is low risk. Generating numbers that go into a signed valuation report is high risk. Apply proportionally heavier review to the latter.
  2. Restrict which tasks AI can touch unsupervised. No AI-generated legal citation, financial figure, or regulatory reference goes to a client without independent verification against a primary source.
  3. Version-lock and log the model. Know exactly which model version produced which output; models get updated and behavior shifts without warning.
  4. Train reviewers on known failure modes, not general AI literacy. A tax reviewer needs to know AI tools fabricate IRS revenue ruling numbers; a valuation reviewer needs to know AI tools misapply comparable-company multiples.
  5. Define the sign-off chain in writing, naming who is accountable at each stage, before the engagement starts, not after a problem surfaces.
  6. Loop in the insurer and general counsel on any new high-risk AI use case before, not after, rollout.

For a broader regulatory grounding, the NIST AI Risk Management Framework (US, voluntary but widely referenced) is a solid free baseline for structuring exactly this kind of guardrail program.

🎬 [VIDEO: "AI and the Law: Liability, Ethics, and Legal Risks" - youtube.com - a walkthrough of real sanctioned cases where generative AI produced fabricated legal citations, and what courts required afterward]

Key Takeaways

  • Liability defaults to the firm and the signing professional; AI vendor disclaimers do not transfer accountability for the final professional opinion.
  • Existing standards of care already apply to AI-assisted work (see ABA Formal Opinion 512, 2024); there is no separate, lighter "AI made the mistake" defense.
  • A defensible human sign-off requires traceability, independent verification against primary sources, and a reviewer competent enough to catch known AI failure modes, not a cursory glance.
  • Keep a documentation trail (tool version, prompt, raw output, edits, reviewer, date) so an incident two years later can be reconstructed and defended.
  • Treat the insurance conversation proactively: insurers are already pricing AI governance into professional indemnity underwriting, and a written policy improves the firm's negotiating position.