# Where AI actually breaks in advisory work
In June 2023, two lawyers filed a federal court brief citing six cases that did not exist. Their research tool, ChatGPT, had invented plausible-sounding precedents, complete with fake docket numbers. A judge in the Southern District of New York fined them and their firm. That case, *Mata v. Avianca*, is now the textbook opening slide in every legal AI training deck. But the underlying failure mode did not stay in legal practice. It is the same failure mode showing up in due diligence memos, valuation models, and audit workpapers across professional services in 2026.
This lesson dissects where AI breaks specifically in advisory deliverables, and what checks catch it before it reaches a client.
Consulting, law, audit, and advisory firms sell judgment, not just output. A manufacturing client doesn't buy a spreadsheet, they buy your firm's conclusion that an acquisition target is worth $340 million and not $410 million. When AI quietly corrupts an input, the error inherits the credibility of the firm's brand. That is different from a retail chatbot mistake: the client has no independent way to check your reasoning, because they hired you precisely because they lack that expertise.
This is sometimes called an information asymmetry problem: the client cannot verify the work as well as the firm can, which is exactly why professional obligations (fiduciary duty, professional negligence standards) exist. AI errors exploit that asymmetry silently.
HallucinationHallucinationA hallucination is when an AI model generates output that is fluent and confident but factually wrong, fabricated, or unsupported by its source data.Voir la définition complète → means a model generates fluent, confident content that is factually false, with no signal to the user that it's fabricated. In advisory work this shows up as:
The danger isn't that the model is wrong sometimes. It's that hallucinated text is stylistically indistinguishable from accurate text. A partner skimming a 40-page draft will not catch a fabricated footnote unless someone actually clicks through to the source.
Check that catches it: mandatory source-verification pass, where every citation, case, or statistic in an AI-assisted draft is independently pulled up and confirmed before the document leaves the building. Some firms now require a "citation audit log" as a deliverable artifact, not just a QA step.
Financial models built with AI assistance (populating DCFDCFDiscounted Cash Flow (DCF) is a valuation method that estimates an asset's value by projecting future cash flows and discounting them to present value using a required rate of return.Voir la définition complète → templates, generating comps, drafting sensitivity tables) fail differently than hallucinationhallucinationA hallucination is when an AI model generates output that is fluent and confident but factually wrong, fabricated, or unsupported by its source data.Voir la définition complète →. The output looks numerically plausible and internally consistent, but rests on a wrong assumption the model never flagged as uncertain.
Example: an AI tool asked to build a discounted cash flowdiscounted cash flowDiscounted Cash Flow (DCF) is a valuation method that estimates an asset's value by projecting future cash flows and discounting them to present value using a required rate of return.Voir la définition complète → model for a target company pulls a "standard" terminal growth rate or discount rate from training data patterns rather than from the client's actual sector and geography. The spreadsheet computes correctly. The formulas are right. The input is silently stale or generic.
Illustrative only:
Terminal growth rate used by AI draft: 3.0% (generic US long-run GDP proxy)
Correct assumption for target's niche market (per analyst research): 1.2%
Enterprise value impact (illustrative, simplified Gordon Growth sensitivity):
EV = FCF / (WACC - g)
At g = 3.0%: EV = $50M / (0.09 - 0.03) = $833M
At g = 1.2%: EV = $50M / (0.09 - 0.012) = $641M
Difference: ~23% valuation swing from one unflagged assumptionThis is model risk: the risk that a model's output is used with confidence it does not deserve, because errors in structure, data, or assumptions aren't visible from the output alone. US bank regulators formalized this concept in SR 11-7 (Federal Reserve guidance on model risk management), and it now informally shapes how professional services firms think about AI-assisted modeling, even outside banking.
Check that catches it: independent assumption review, where a second person (not the AI, not the original preparer) states every material assumption in plain language and sources it before the model is finalized.
AI tools increasingly triage due diligence: sorting contracts by risk, flagging "red flag" clauses, ranking candidate acquisition targets, or scoring management teams from LinkedIn and public data. When these tools are trained on historical deal data, they can encode past bias into supposedly objective scores.
Concrete pattern: a screening tool trained on past "successful" portfolio companies may implicitly favor management teams that resemble prior (often demographically narrow) leadership, penalizing equally qualified but differently composed teams. This is not hypothetical, it mirrors well-documented hiring-algorithm bias cases, such as Amazon's scrapped AI recruiting tool that downgraded resumes mentioning women's colleges or activities.
In due diligence, biased screening creates both an ethical exposure and a commercial one: you may be systematically deprioritizing good targets or good hires and never know it, because the tool never explains its ranking logic in terms a partner would recognize as biased.
Check that catches it: disparate impact testing on screening outputs before they influence a recommendation, plus a rule that AI-generated rankings are advisory inputs only, never the final filter, with a human required to review anything the tool excludes, not just what it includes.
Vérification des acquis
1. Why does the 'information asymmetry problem' make AI errors especially dangerous in professional services advisory work?
2. What is the defining characteristic of 'hallucination' as a model failure mode, as distinguished from other kinds of AI mistakes?
3. A manufacturing client hires a consulting firm to value an acquisition target. According to the lesson, what makes an AI-driven error in this context different from an AI error in a retail chatbot?
4. Select ALL correct answers about how the hallucination failure mode manifests in advisory deliverables.
Sélectionnez toutes les réponses correctes.
5. Select ALL correct answers about why professional services is described as a 'special risk case' for AI errors.
Sélectionnez toutes les réponses correctes.
Each failure mode above is bad alone. They compound because professional services deliverables move fast through a chain of trust: analyst to manager to partner to client to the client's board. At each handoff, the assumption is "the previous person checked this." AI-generated content breaks that chain because it can pass a fluency check (does it read well?) while failing an accuracy check (is it true?), and most reviewers under deadline pressure only run the first check.
Regulators are starting to respond. The EU's AI Act (entered into force 2024, obligations phasing in through 2026-2027) classifies certain AI uses in professional contexts, including some credit and employment-related screening, as "high-risk," triggering mandatory risk management, documentation, and human oversight requirements. In the US, there is no single federal AI law as of 2026; instead, sector regulators (SEC for investment advisers, state bar associations for legal practice) issue guidance and enforce existing professional conduct rules against AI-caused errors. The FTC (Federal Trade Commission) has also signaled it will treat AI-related misrepresentation to clients as a deceptive practice under existing consumer protection authority.
Before any AI tool touches a client deliverable, professional services firms should have answers to:
1. Provenance: can every factual claim or citation be traced to a verifiable source, independent of the AI tool?
2. Assumption disclosure: does the tool surface (not bury) the key assumptions driving a financial output?
3. Bias testing: has any screening or ranking tool been tested for disparate outcomes across protected characteristics, on a representative sample, before go-live?
4. Human sign-off: is there a named individual, not "the team," accountable for verifying AI-assisted sections before client delivery?
5. Audit trail: can the firm reconstruct, months later, what the AI generated versus what a human edited, if a regulator or client disputes the work?