Request Demo
← All posts

What Counts as a Model Now? Building the AI Inventory

TL;DR

Model risk management guidance has always defined a model broadly — a quantitative method that processes inputs into estimates or decisions — and LLM-based applications fit the definition comfortably, qualitative outputs and all. Institutions increasingly treat them as in scope. The genuinely hard problem is the inventory: AI no longer arrives as a system you commissioned, but embedded in vendor tools and unofficially through staff use. A practical sweep combines procurement and expense review, vendor-feature audits, staff attestation, and network telemetry. Tier what you find by decision impact — customer-outcome-affecting systems owe full validation; internal drafting tools owe a usage policy — and remember that the most expensive system in the inventory is the one that is not in it.

The definition was never the hard part

Ask whether a large language model application “counts as a model,” and the honest answer is that the question was settled before LLMs existed. Model risk management guidance (SR 11-7 and its kin) defines a model broadly and functionally: a method that applies theories, techniques, and assumptions to process input data into estimates or decisions. Nothing in that definition requires a regression, a score, or even a number. A system that reads a dispute file and produces a recommended determination is processing inputs into a decision by way of quantitative technique — several billion parameters of it. The occasional objection that LLM outputs are “qualitative” misreads the guidance twice: the definition never excluded qualitative outputs, and supervisory practice has long treated qualitative-adjustment tools as in scope. Institutions have largely stopped litigating the point; the prevailing posture is to treat LLM-based applications as models — or, at minimum, to document explicitly why a given one is not. Either way, something must be written down, which is the real subject of this post.

AI does not arrive through the front door

The classical inventory assumed models arrive by commissioning: someone built or bought a thing called a model, and it was entered on a list. AI arrives differently, along two quieter routes. It arrives embedded: the vendor platform you licensed for case management ships an LLM summarization feature in a quarterly release; the productivity suite adds a drafting assistant; a browser extension answers questions about whatever page is open, including the page displaying customer data. No procurement event marked any of these as a model acquisition. And it arrives unofficially: an analyst pastes a fact pattern into a public chatbot to get a first draft, because the draft is good and the queue is long. Shadow use is not exotic misconduct — it is diligent people reaching for the best available tool. The inventory that catches only commissioned systems will miss most of what is actually running.

A practical sweep

No single instrument finds everything, so the workable approach layers four, each described here at the level of method rather than product.

The sweep is not a one-time census. Embedded features ship quarterly; the inventory that was accurate in March is a historical document by June. Institutions increasingly treat the sweep as a recurring control with an owner, not a project.

Tier by decision impact

An inventory that treats every discovered system identically will collapse under its own weight, so the second move is tiering — and the cleanest axis is decision impact. At the top: systems whose outputs affect customer outcomes — dispute determinations, credit decisions, alert dispositions, complaint responses. These owe the full apparatus the guidance describes: documented purpose and limitations, validation appropriate to the technology (a genuinely open problem for LLMs, which we have written about separately), ongoing monitoring, and effective challenge. Below that: internal drafting and summarization tools whose outputs a human revises before anything depends on them. These owe a usage policy — what data may enter them, what outputs require review, what uses are prohibited — plus periodic confirmation that actual use matches the policy. The tier boundary deserves care, because drift across it is silent: a summarization tool becomes decision-affecting the day its summaries start being the only version of the file anyone reads. Tiering is a judgment the institution should be prepared to explain, not a label applied once.

The system nobody listed

Why does the sweep deserve this much effort? Because of how the un-inventoried system surfaces. It does not announce itself; it appears mid-examination, when a reviewer asks how a particular determination was reached and the trail runs through a tool that is on no list, under no policy, with no validation record and no owner. At that moment the institution has two findings instead of one: whatever the tool did, and the demonstrated gap in the inventory control itself — the second being worse, because it puts every other assurance in doubt. The examination is the user interface through which your governance is actually experienced, and an inventory is the index page of that interface. The question is not whether AI is in the building — it is. The question is whether you found it before someone else did.