AIAI in BankingBankingFintech

SR 11-7 still bites, and gradient boosting just made the wound worse

SR 11-7 was written for logistic regression, but banks are now deploying gradient boosting, neural networks, and foundation model-powered scoring into production. This piece unpacks what explainability actually means under the guidance, why examiners are pushing harder on it in 2026, and where the governance frameworks genuinely break down.

Neo NeumannNeo NeumannAI Practice LeadSeptember 25, 2026

Listen to the podcast

4 min

Chapters

Key takeaways

  • Pull ten recent denials from your most complex model and try to write each adverse action reason by hand from the model's own outputs.
  • Treat any model where you cannot reconstruct an individual adverse decision as unvalidated, whatever the documentation claims.
  • Do not present SHAP charts as a full explanation; they show what the model did, not whether the decision was defensible.
  • Avoid surrogate explainer models, since they can match the real model 95% of the time and diverge on the denials that carry legal risk.
  • Price gradient boosting's few points of predictive lift against a revalidation cost you pay at every refresh and retrain.
Read the full transcript

Host:Leaders' insights. Five minutes on SR11, seven still bites, and gradient boosting just made the wound worse. If a bank can't explain why its model denied someone a mortgage, should regulators just shut the model down?

Expert:That's exactly the fight happening right now. SR11-7. The Federal Reserve's model risk guidance from 2011 doesn't literally say shut it down. But examiners in 2026 are treating unexplainable models as unvalidated, and an unvalidated model can't go into production. So functionally, yes.

Host:15-year-old guidance. Why is it suddenly biting harder this year?

Expert:Because the models changed and the rules didn't. SR11-7 was written when everyone scored credit with logistic regression, a simple statistical method where you can read off exactly how much each factor moves the outcome. Now banks are shipping gradient boosting, which stacks hundreds of little decision trees on top of each other, and the interpretability just evaporates.

Host:Give me the moment this became a headline.

Expert:The trigger was examiners flagging Foundation model-powered scoring, big pre-trained AI exams repurposed for credit decisions in supervisory exams earlier this year. Google's AI team has claimed their explainability tooling can attribute a score down to individual features with high fidelity, but they sell that tooling, so I'd cross-check that against a validation team that doesn't have a product to move.

Host:Let me throw three things people believe at you. First, SHAP values solve the explainability problem. True, half true, or wrong?

Expert:Half true, and the dangerous half. SHAP, a method that assigns each input a contribution to the final score, gives you a number for every feature. The problem is, those numbers describe correlation inside the model, not the causal story a regulator wants. You can hand an examiner a beautiful SHAP chart, and they'll ask, so would approving this person have been the right call? And SHAP has nothing to say about that.

Host:So it explains the math, not the decision.

Expert:Right, it tells you what the model did, not whether what it did was defensible. Under SR11, 7, that gap is where validations die.

Host:Second belief, gradient boosting is inherently more accurate, so the explainability tradeoff is worth it. Where does that land?

Expert:Mostly wrong for credit, actually. On tabular data, rows and columns, income, balances, payment history, gradient boosting often beats logistic regression by a few points of predictive lift. But a few points on a portfolio where you already have decades of clean data rarely justifies losing the ability to answer a fair lending complaint. OpenAI has published numbers suggesting large models close reasoning gaps fast. Though again, they're selling the models. So treat the benchmark like marketing until someone independent reruns it. You're saying the accuracy gain is real but small,

Host:and the governance cost is large.

Expert:And the governance cost is fixed, not one time. Every model refresh, every retrain, you revalidate. A logistic model you validate once and barely touch. A boosted model drifts, so you're paying that tax quarterly.

Host:Third, banks can just bolt on a simpler explainer model next to the complex one and satisfy the regulator.

Expert:Wrong. And examiners caught on to this one specifically. The trick is called a surrogate. You train a simple, readable model to mimic the complex one's outputs, then show the regulator the readable version. The issue is the surrogate can agree with the real model 95% of the time and disagree exactly on the edge cases, the denials, where the legal risk lives. You've built an explanation that's most wrong precisely

Host:when it matters. So the honest answer to my opening question?

Expert:If you can't reconstruct the reason for an individual adverse decision, not the average, the individual, you shouldn't have that model in front of consumers. That's not a regulator being difficult. That's the actual law behind adverse action notices, the letter you're owed explaining why you got denied.

Host:Leave the risk officer listening with one thing to do, Mundy.

Expert:Pull ten recent denials from your most complex model and try to write the adverse action reason for each by hand from the model's own outputs. If you can't do it for even one, that model isn't validated no matter what your documentation says. And it's better you find that out before the examiner does.

Host:What we read for this one, TechCrunch AI, ours Technica AI, the decoder, KD Nuggets, Google AI, vendor, AI Lab, OpenAI, vendor, AI Lab. Done for today. There's a new AI piece every morning at MBA-training.com.

The concept at stake is model explainability as a compliance obligation, not a technical preference. Banks have been building explainability tooling for years, mostly as a way to satisfy internal model risk committees and, on occasion, satisfy an adverse action notice requirement under the Equal Credit Opportunity Act. What has changed is the intensity of examiner scrutiny, the complexity of the models now in production, and the uncomfortable gap between what SHAP values tell a reviewer and what the model is actually doing.

SR 11-7, the Federal Reserve and OCC guidance issued in 2011, predates transformer architectures, gradient boosting frameworks, and the current generation of LLM-powered credit tools by a decade. But it has not been replaced. Examiners are applying it to models its authors never imagined.

How model risk flows into provisions, capital and consent orders

A bank's credit portfolio is priced on assumptions about default probability. Those assumptions live inside models. When a model misbehaves, the loss does not stay in the model risk management (MRM) team's inbox. It flows into provision charges, capital requirements under Basel III, and, if the model has been making discriminatory decisions, into consent orders with the CFPB or DOJ.

The OCC's 2023 model risk management handbook update made clear that "conceptual soundness" requires being able to explain not just what a model predicts but why it predicts it, and whether that reasoning is consistent with the bank's intended use. For a gradient boosting model doing mortgage underwriting, "consistent with intended use" is a very demanding standard when the model has learned from historical data that reflects decades of discriminatory lending patterns.

Regulators are also watching how banks deploy third-party models. JPMorgan Chase, Wells Fargo, and Bank of America all use vendor-supplied scoring models for segments of their credit decisions. Under SR 11-7, the bank owns the model risk regardless of who built it. If the vendor cannot produce a model fact sheet that satisfies an examiner, the bank answers for it.

Beyond credit,catching model risk before it triggers a supervisory action requires banks to apply the same explainability logic to AML transaction monitoring, fraud detection, and increasingly to the LLM-powered tools entering compliance workflows. An LLM that surfaces suspicious activity alerts but cannot articulate its reasoning chain is, under SR 11-7's logic, an unexplained model producing outputs the bank is relying on.

What does SR 11-7 require for model explainability?

SR 11-7 does not use the word "explainability." It requires conceptual soundness, ongoing monitoring, and documentation sufficient for a "knowledgeable party" to evaluate the model. In practice, MRM teams have interpreted this to mean three things:

First, the model must have a documented conceptual narrative: a plain-language account of why the chosen approach is appropriate for the intended use, what assumptions it rests on, and what it would take to invalidate those assumptions. For a logistic regression, this is straightforward. For an XGBoost model with 400 features, it requires discipline.

Second, individual predictions must be explainable at the decision level, particularly for adverse actions. The Fair Housing Act and ECOA require that applicants denied credit receive specific reasons. Post-hoc explanation tools like SHAP (SHapley Additive exPlanations) assign contribution scores to each feature for each prediction, giving the bank a basis for generating those reasons. But SHAP explanations approximate the model's behavior; they do not describe it exactly. When a CFPB examiner asks whether the reason codes produced are accurate, the honest answer is "approximately, within defined tolerance bounds," and the bank needs documentation to show those bounds have been validated.

Third, monitoring must detect when the explanation structure itself has shifted. If a model was approved with home value as its second-highest SHAP contributor and, six months later, a geographic cluster feature has displaced it, that is a material change that should trigger a model review. Many banks run this monitoring quarterly. Given how fast data distributions can shift, that cadence is increasingly inadequate.

A concrete example: in 2024, a mid-sized regional bank in the Southeast discovered during a fair-lending examination that its auto loan pricing model, built on LightGBM, was using postal code as a top-five feature. The model had been approved internally on the basis that postal code was a proxy for dealership network costs. The examiner read it differently: as a proxy for race in several metro markets. The bank's SHAP documentation showed the feature's contribution but had no analysis of its disparate impact. The result was a memorandum of understanding requiring model remediation and a second-look program for affected applicants. The cost ran past $40 million including remediation, monitoring infrastructure, and legal fees.

Where SHAP holds up and where it fails on LLMs

Post-hoc explainability tools work well when the model is relatively stable, the feature set is interpretable by domain experts, and the prediction task has a clear ground truth that can be validated over time. Traditional credit scoring, fraud rule optimization, and collateral valuation models sit in this zone.

They work poorly for foundation models and LLMs being used in credit or compliance workflows. When a bank deploys an LLM to summarize a commercial borrower's financial statements and produce a risk narrative, the explanation problem is qualitatively different. There is no feature importance vector. The "reasoning" is a function of attention patterns across billions of parameters. Asking SHAP to explain an LLM's output is roughly like asking a flight data recorder to explain the pilot's intention.

Building credit models that survive fair-lending scrutiny increasingly means choosing model architectures with governance in mind from the start, not retrofitting explainability after deployment. Constrained models, monotonic gradient boosting, or scorecard ensembles with explicit feature ceilings are less accurate than unconstrained deep models on holdout sets. They are also far more defensible when an OCC examiner sits down with your MRM documentation.

The honest tradeoff: a bank that optimizes purely for predictive lift will deploy models that perform better on paper and survive examinations worse. The margin difference between an 8% Gini improvement and a consent order is not close.

SR 11-7 will eventually be superseded by more specific AI guidance, and the OCC and Fed have signaled updates are in progress. Until that happens, the 2011 text remains the operational standard, and the banks treating explainability as a documentation exercise rather than a design constraint are building a liability, one model deployment at a time.

The full course on this sector:AI in Banking.

Frequently asked questions

Does SR 11-7 actually mention explainability?

No. SR 11-7, issued in 2011 by the Federal Reserve and OCC, requires conceptual soundness, ongoing monitoring and documentation sufficient for a knowledgeable party to evaluate a model. Model risk management teams have translated that into three demands: a documented conceptual narrative, decision-level explanations for adverse actions, and monitoring that catches shifts in the explanation structure itself.

Are SHAP values enough to satisfy a CFPB examiner?

Not on their own. SHAP explanations approximate a model's behaviour rather than describe it exactly, so a bank must document validated tolerance bounds for the reason codes it produces. Feature contribution scores also say nothing about disparate impact, which is what fair-lending examiners look for.

Who owns model risk when a bank buys a vendor scoring model?

The bank does. Under SR 11-7, model risk stays with the institution regardless of who built the model, and JPMorgan Chase, Wells Fargo and Bank of America all run vendor-supplied scoring in parts of their credit decisions. If the vendor cannot produce a model fact sheet an examiner accepts, the bank answers for it.

What happens when a credit model uses postal code as a feature?

It can be read as a proxy for race. In 2024 a mid-sized Southeastern regional bank found during a fair-lending examination that its LightGBM auto loan pricing model ranked postal code in its top five features, justified internally as a proxy for dealership network costs. The outcome was a memorandum of understanding, a second-look program and costs past $40 million.

Go deeper

The lessons that take this article further, free to read.

  1. 1Spotting model risk before it becomes a loss eventAI in banking
  2. 2Building credit decisioning models that survive fair-lending scrutinyAI in banking
  3. 3How the regulatory map for AI in banking actually fits togetherAI in banking
  4. 4Running the pre-deployment gauntlet: checks that catch problems earlyAI in banking
  5. 5Building a compliance function that survives an auditBanking: how the sector works

Finished reading?

Validate your read to earn XP and feed your radar.