Request Demo

Research & announcements

Protocol releases, technical thinking, and updates from Arcus Labs

Occasional notes on new research and releases — a couple a month at most. We’ll send a confirmation link; unsubscribe anytime from any email with one click.

Almost there — check your inbox

We sent a confirmation link to finish your subscription. Click it and you’re all set.

Anatomy of a Regulator Dossier
Anatomy of a Regulator Dossier

Imagine the first day of an examination, and one automated decision under review. What would a complete record look like — the document you could hand over the table without a caveat? Here is that record as a reference architecture: nine parts, from provenance and tamper-evidence through reasoning, citations, limitations, and attestation, with every line color-coded by what kind of thing it is.

Read →
What Counts as a Model Now? Building the AI Inventory
What Counts as a Model Now? Building the AI Inventory

Model risk management guidance defines a model broadly — a method that processes inputs into estimates or decisions — and LLM-based applications fit. The hard part is not the definition; it is the inventory. AI arrives embedded in vendor tools and unofficially through staff use, and the system nobody listed is the one that surfaces during an examination.

Read →
Claim, Grounds, Warrant: Toulmin's The Uses of Argument
Claim, Grounds, Warrant: Toulmin's The Uses of Argument

In 1958, Stephen Toulmin argued that real arguments do not run on syllogisms: they run on a claim, the grounds beneath it, and a warrant licensing the step between — plus the qualifier and rebuttal that honest conclusions carry. An adverse-action letter is a Toulmin layout with the hardest slots left blank.

Read →
The First-Day Letter, AI Edition
The First-Day Letter, AI Edition

Before the examination team arrives, the institution receives a document request — the first-day letter. AI-assisted decisioning now belongs in it. Here is a draft of that section: eight numbered items an institution should be able to produce same-day, and why the production clock starts running before any individual decision is reviewed.

Read →
The Most Important Answer Is 'I Don't Know': Reading AbstentionBench
The Most Important Answer Is 'I Don't Know': Reading AbstentionBench

Some questions have no answer — the premise is false, the facts are missing, the request is underspecified. A 2025 benchmark asked whether frontier models can say so. They cannot, reliably — and models fine-tuned for reasoning often abstain less, answering confidently where no answer exists. In regulated work, that is the exact behavior a decision system must be built to refuse.

Read →
You Don't Need Smarter Models
You Don't Need Smarter Models

A common plan hiding inside AI roadmaps: wait — the next generation will be good enough to trust. But capability and defensibility are different axes. Scaling moves one and leaves the other exactly where it was; a more capable model in an ungoverned harness produces the same indefensible decision, better written.

Read →