The First-Day Letter, AI Edition
Before a supervisory examination begins, the institution receives a document request — the first-day letter — listing the records the examination team will work from. AI-assisted decisioning now belongs in that letter, and this post drafts the section: eight numbered items, from the system inventory to the reproducibility protocol. Institutions can read it as a self-test — could you produce these same-day? — and supervision teams as a starting draft. The framing point comes before any individual item: whether production takes an afternoon or three weeks is itself the first finding of the examination.
The request that precedes the questions
Every examination begins the same way: before anyone sits down, the institution receives a request for documents. The list is rarely exotic — policies, committee minutes, management reports, a sample of files — but it is never neutral. What the letter asks for defines what the examination considers examinable, and institutions read it that way: the first-day letter is the clearest signal an examined entity ever gets about what its supervisor believes good governance produces.
As AI-assisted systems move from drafting assistance into decisioning — disputes, adverse action, claims, alert triage — the letter needs a section it historically did not have. Below is a draft of that section, written to be useful in both directions. An institution can treat it as a preview and a self-test. A supervision team can treat it as a starting point. And it is deliberately vendor-neutral: nothing in it assumes a particular product or architecture, only that decisions were made and the institution can account for them.
The eight items
One reading note: “sampled decisions” below means a sample the examination team selects — not one the institution curates. With that understood, the request:
- System inventory. A current inventory of AI-assisted decision systems, each with a named owner, a tier reflecting decision impact, and the rationale for what was scoped in and out. Good production reads like a maintained register, not a reconstruction — and the out-of-scope rationale gets read as closely as the list.
- Reasoning records for sampled decisions. The complete reasoning record for each sampled decision — the steps taken, the evidence cited at each step, and the outcome — as it existed at decision time. Good production is retrieval, not reassembly: the record was sealed when the decision issued, and nothing in it postdates the decision.
- Validation and effective-challenge reports. What was tested, at what granularity — the step level, not only the outcome level — by whom, and how the challengers were independent of the build. Good production names the tests and the testers, and shows challenge that changed something.
- Change history. Model versions, prompt and rule-corpus versions, and configuration changes, each mapped to the date ranges of the decisions it governed. Good production answers “which system, exactly, made this decision in March” without a meeting.
- Human-oversight evidence. Reviewer actions on the sampled decisions, override and approval rates, and how the institution verifies that reviewer attention is meaningful rather than pro forma. A 100% approval rate with a median review time of eleven seconds is production too — of a different finding.
- Exception and incident log. Exceptions, escalations, and incidents involving AI-assisted decisions, with dispositions. Good production shows a log that is used: entries with owners, resolutions, and dates — not an empty table offered as evidence of health.
- Vendor documentation and diligence. For any third-party AI embedded in decision flows, the vendor’s documentation plus the institution’s own diligence over it. The institution owns the decision even where it licensed the reasoning, and the diligence file is where that ownership is visible.
- Reproducibility protocol. The documented procedure for re-running a past decision: who may invoke it, what is held fixed, and what “matching” means. Good production is a protocol that has actually been executed, with results on file — not a capability asserted in a policy document.
None of these items asks whether the AI was right. That comes later, decision by decision, from the records item two produces. The letter asks something prior: whether the institution can account for its own decisioning at all.
Production time is the first finding
Here is the punchline, and it lands before a single decision is reviewed: the difference between same-day production and a three-week scramble is the examination finding. Every item above is something a well-governed operation produces as a byproduct of running. If the records exist because the system created them at decision time, production is a retrieval job. If they must be assembled — logs correlated after the fact, versions reconstructed from deployment tickets, narratives re-derived from whatever survived — production is a project, and the length of that project measures the gap between how decisions were made and how they can be accounted for.
Model risk management guidance (SR 11-7) has always treated documentation this way: as evidence of control, not paperwork adjacent to it. Architectures that seal the reasoning record at decision time reduce and interrupt the failure modes that make production slow — reconstruction, version ambiguity, the post-hoc narrative — though nothing removes the obligation to check. The examination still happens. It just starts from evidence instead of testimony.
A first-day letter is a mirror held up to an institution’s governance: it does not create the record, it reveals whether one exists. Institutions that can answer this section same-day will not find the letter demanding — because they were, in effect, answering it every day the system ran.