Imagine the first day of an examination, and one automated decision under review. What would a complete record look like — the document you could hand over the table without a caveat? Here is that record as a reference architecture: nine parts, from provenance and tamper-evidence through reasoning, citations, limitations, and attestation, with every line color-coded by what kind of thing it is.
Read →Research & announcements
Protocol releases, technical thinking, and updates from Arcus Labs
Almost there — check your inbox
We sent a confirmation link to finish your subscription. Click it and you’re all set.
Model risk management guidance defines a model broadly — a method that processes inputs into estimates or decisions — and LLM-based applications fit. The hard part is not the definition; it is the inventory. AI arrives embedded in vendor tools and unofficially through staff use, and the system nobody listed is the one that surfaces during an examination.
Read →
In 1958, Stephen Toulmin argued that real arguments do not run on syllogisms: they run on a claim, the grounds beneath it, and a warrant licensing the step between — plus the qualifier and rebuttal that honest conclusions carry. An adverse-action letter is a Toulmin layout with the hardest slots left blank.
Read →
Before the examination team arrives, the institution receives a document request — the first-day letter. AI-assisted decisioning now belongs in it. Here is a draft of that section: eight numbered items an institution should be able to produce same-day, and why the production clock starts running before any individual decision is reviewed.
Read →
Some questions have no answer — the premise is false, the facts are missing, the request is underspecified. A 2025 benchmark asked whether frontier models can say so. They cannot, reliably — and models fine-tuned for reasoning often abstain less, answering confidently where no answer exists. In regulated work, that is the exact behavior a decision system must be built to refuse.
Read →
A common plan hiding inside AI roadmaps: wait — the next generation will be good enough to trust. But capability and defensibility are different axes. Scaling moves one and leaves the other exactly where it was; a more capable model in an ungoverned harness produces the same indefensible decision, better written.
Read →