Vendor Diligence for “AI-Powered” Compliance Tools
Third-party risk guidance is blunt on one point: outsourcing an activity does not outsource the accountability. When the tool is “AI-powered,” that principle turns into a specific demand list — per-decision reasoning records as a contractual right, version pinning with change notification, evaluation evidence at the step level rather than headline accuracy, data provenance and data-use terms, disclosure of sub-processors and upstream model dependencies, examination support obligations, and exit terms that export your decision records in usable form. And beware the analogical trap: “it’s like the system you already validated” is a claim about surface similarity until the vendor can show the properties the prior validation actually rested on.
The phrase that answers nothing
“AI-powered” is a statement about the vendor’s architecture. Your obligations, though, attach to your decisions — supervisory guidance on third-party relationships has said for years that an institution’s accountability for an activity is not diminished by outsourcing it, and institutions increasingly read that to cover the reasoning inside a vendor’s tool, not just its uptime and its security posture. So the purpose of AI vendor diligence is concrete: when a specific determination is challenged — by an examiner, an auditor, or a customer’s attorney — can you reconstruct and defend it? Every question below is that one question, decomposed. A vendor’s comfort answering them is itself diagnostic: governable tools were built expecting these questions.
The demand list
- Per-decision reasoning records, as a contractual right. For each decision the tool touches, you need the record of how it was reached — inputs considered, steps taken, evidence cited — accessible to you on your schedule. Contractually, not as a favor: a record available at the vendor’s discretion, through a support ticket, at the vendor’s pace, is not a record you control. If the answer is “we log prompts and outputs,” note what is missing: the reasoning between them.
- Version pinning and change notification. Model, prompts, rules, retrieval corpus — each is configuration, and each changes. You need to know which configuration produced which decision, which means versions with dates, notice before material changes, and ideally the ability to pin. A tool whose behavior shifted silently in April cannot explain its March decisions.
- Evaluation evidence at the step level. Headline accuracy on a validation set tells you how often the tool lands on defensible outcomes, not whether its reasoning would survive review — a right answer reached through a wrong step is a liability wearing a success. Ask what the vendor measures per step: citation validity, evidence grounding, consistency across identical fact patterns, abstention on thin records. If the steps are not measured, ask whether they exist as measurable things at all.
- Data provenance — in both directions. What was the underlying model trained or fine-tuned on, at least in kind? And what may the vendor do with your data: does it leave your tenancy, train anyone’s models, persist after processing? “We may use customer data to improve our services” is a sentence to read slowly.
- Sub-processors and upstream model dependencies. Many “AI-powered” tools are an interface over someone else’s model. You need to know whose, under what terms, and with what notice when it changes — because your tool’s behavior can change when a company you have no contract with ships an update.
- Examination support obligations. When your examiner asks about the tool, what has the vendor committed to provide, in what form, in what timeframe? Documentation suitable for supervisory review is a deliverable to specify in the contract, not an assumption to discover under deadline.
- Exit terms. Decisions outlive subscriptions. On termination, your decision records — including the reasoning records above — must export in a form usable without the vendor’s software. An archive readable only through the product you just left is retention in name only.
The analogical trap
One reassurance deserves its own section, because it is the most seductive: “this works like the system you already validated.” The rules-based screener, the previous scoring model, the workflow tool the examiners have seen for years — the new AI product is presented as that thing, upgraded. The reasoning move here is analogy, and analogy’s known failure mode applies with full force: a mapping built on surface features — same inputs, same outputs, same screen layout — borrows confidence from a system whose deep structure was entirely different. Your prior validation rested on specific properties: the rules were enumerable, behavior was deterministic, identical inputs produced identical outputs, changes required a release you approved. The analogy transfers trust legitimately only if those load-bearing properties carried over — so name them, and ask the vendor to show each one, or to say plainly which ones no longer hold and what replaces them. A vendor who can answer has earned the comparison. A vendor who repeats the analogy louder has told you it is surface.
What a good answer sounds like
This list will strike some vendors as burdensome, and that reaction is data. The vendors building for regulated institutions are converging on the same properties from their side — decision-level records, pinned configurations, step-level metrics, examiner-ready exports — because effective challenge needs a surface to act on, and a black box offers none. None of this makes a tool safe; diligence reduces and interrupts the ways vendor opacity becomes your finding, it does not remove them. But the asymmetry at the table is real: you are not asking for favors. You are asking whether the tool can participate in your governance. Tools that can were built to answer. Tools that cannot were built to be trusted — which is precisely the thing you are not permitted to do on someone else’s word.