Request Demo
← All posts

The Agencies Carved Gen-AI Out of SR 26-2. So Who Governs It?

TL;DR

The April 2026 interagency model risk guidance — SR 26-2 — modernized fifteen years of model risk management, and in the process drew a boundary that has received almost no attention: generative-AI and agentic-AI models are out of the guidance’s scope. That is not an exemption from oversight. It means the fastest-growing category of models in banking now has no tailored supervisory framework at all — while safety-and-soundness expectations, consumer protection law, and examiner curiosity all still apply. Institutions deploying LLMs face a choice: wait years for gen-AI guidance, or extend model-risk discipline to gen-AI voluntarily, on architecture the eventual guidance will almost certainly require anyway.

The sentence nobody is quoting

Most coverage of SR 26-2 has focused on what changed: prescriptive requirements out, risk-based principles in; rigid tiering out, materiality-scaled oversight in. Fair enough — those are real shifts, and we have written about them. But the guidance also does something quieter. It draws its own boundary, and generative and agentic AI sit outside it. The agencies took fifteen years of supervisory experience with statistical models and declined to extend the framework to the model class that every bank’s innovation team is deploying right now.

Read charitably, the carve-out is honest. The MRM framework’s core machinery — conceptual soundness review, outcomes analysis, benchmarking against holdout data — was built for models whose mechanics can be inspected and whose behavior is stable across a defined input space. As we argued when SR 11-7 still governed, LLMs break those assumptions structurally. Rather than pretend the old validation toolkit maps cleanly onto generative systems, the agencies scoped them out. The intellectual honesty is admirable. The practical result is a vacuum.

Out of scope is not out of trouble

Here is what the carve-out does not do. It does not repeal safety-and-soundness authority, which reaches any practice that threatens an institution’s condition, framework or no framework. It does not touch consumer protection law — ECOA and fair lending, UDAP/UDAAP, Reg E and Reg Z obligations all apply to decisions regardless of what kind of system made them. It does not bind state regulators, or the CFPB, or the plaintiffs’ bar. And it does not stop an examiner from asking the only question that ever mattered: show me how this system arrived at this decision. The carve-out removes the map, not the territory.

That leaves LLM deployments in the worst quadrant: full exposure, no tailored framework. Under SR 11-7, at least, banks knew what documentation and validation examiners expected and could point to compliance with it as evidence of prudence. A bank running LLMs in consequential workflows today cannot point to anything — there is no safe harbor to comply with. When something goes wrong, the institution will be judged against general prudential expectations, retroactively, by examiners with maximal hindsight and zero obligation to accept “the guidance didn’t cover it” as a defense. Ungoverned space is not lenient space. It is unpredictable space, and banks price unpredictability as risk.

What fills a vacuum

Three things, historically. First, voluntary extension: prudent institutions apply the spirit of the nearest framework to the uncovered activity, and their practices become the de facto standard examiners test everyone else against. Second, enforcement: a public failure produces consent orders, and the industry reverse-engineers expectations from the penalties. Third, eventual guidance: the agencies watch both, then codify. Every institution deploying gen-AI today is choosing, knowingly or not, which of these to bet on. Betting on enforcement is betting your institution is not the case study. Betting on eventual guidance is betting your deployments stay perimeter-grade until it lands. The only bet with a good expected value is the first one.

And voluntary extension is more tractable than it sounds, because SR 26-2’s principles — as opposed to its legacy validation mechanics — translate to generative systems surprisingly well. Materiality-scaled oversight translates directly: an LLM drafting marketing copy and an LLM informing adverse action deserve different rigor. Effective challenge translates directly: someone qualified, incentivized, and empowered must genuinely challenge the system’s outputs and design. The three components of validation translate at the level of intent: conceptual soundness becomes is the reasoning process sound and appropriate for the use; outcomes analysis becomes are results systematically evaluated against ground truth where it exists, including fairness outcomes; ongoing monitoring becomes is behavior tracked across prompt changes, model versions, and drift. What does not translate is the evidence base. Coefficients and holdout curves are gone. The only artifact that can take their place is the reasoning record itself.

The architecture the eventual guidance will require

This is a prediction, and we are comfortable making it: when gen-AI guidance eventually arrives, its evidentiary core will be reasoning-level documentation — because there is nothing else it could be. Every precedent points the same way. SR 26-2 kept effective challenge and process documentation at its center. The EU AI Act built its high-risk regime around traceability and logging. Colorado requires documentation of the decision process, not just the decision. Regulators converge on the same demand because it is the only demand that works for systems without inspectable parameters: make the reasoning inspectable instead.

An institution that runs its generative systems on externalized, structured reasoning — explicit steps, verification gates, documented challenge, sealed traces — is compliant with the spirit of the current framework and pre-adapted to any plausible future one. It can answer the examiner’s question today, under general prudential authority, and it will be able to answer it under whatever letter the agencies publish next. That is what we build IRG for. But vendor aside, the strategic point stands: the carve-out is not a holiday from governance. It is an interval in which the institutions that take reasoning governance seriously get to define what adequate looks like — and the ones that don’t get to be measured against them.

SR 26-2 didn’t exempt generative AI. It admitted the old tools can’t govern it — and left the industry to build the ones that can. The banks that build first will set the standard everyone else is examined against.