Request Demo

Research & announcements

Protocol releases, technical thinking, and updates from Arcus Labs

Occasional notes on new research and releases — a couple a month at most. We’ll send a confirmation link; unsubscribe anytime from any email with one click.

Almost there — check your inbox

We sent a confirmation link to finish your subscription. Click it and you’re all set.

The Examiner Is the User
The Examiner Is the User

Every AI system in a regulated workflow has two users: the operator who requests the output today, and the examiner who reviews it later — with subpoena-grade patience and no goodwill. Almost all AI product design serves the first user. The systems that reach core workflows will be the ones designed for the second.

Read
Artificial Intelligence Is Not Artificial Reasoning
Artificial Intelligence Is Not Artificial Reasoning

The industry uses one word for two different things: the capability and the process. AI asks a model what it thinks; artificial reasoning defines how the system must think — as an explicit computation that records itself as it runs. Here is the distinction in full, what the term does not claim, and the falsifiable bet underneath it.

Read
Why the Numbers Should Never Come From the Model
Why the Numbers Should Never Come From the Model

Every number an LLM produces is a prediction of what a number would look like. That's tolerable in a chatbot and disqualifying in a decision system. The engineering answer is deterministic tool-nodes — zero tokens, byte-reproducible — but the property that actually matters is authority: the model may cite the numbers, never recompute or override them.

Read
The 10/45/90-Day Problem: Reg E Error Resolution at Understaffed Scale
The 10/45/90-Day Problem: Reg E Error Resolution at Understaffed Scale

Reg E’s error-resolution deadlines are absolute: 10 business days to investigate, 45 or 90 calendar days with provisional credit, no tolling for backlogs or turnover. Dispute volume has exploded with P2P fraud while investigation teams haven’t grown. The failure mode isn’t missing deadlines — it’s making them with investigations that can’t survive an exam.

Read
The Four-Fifths Rule, Explained for AI Decisioning
The Four-Fifths Rule, Explained for AI Decisioning

The four-fifths rule is a fifty-year-old screening heuristic that AI decisioning has made load-bearing again. It is a tripwire, not a verdict — and both halves of that sentence matter. What the Adverse Impact Ratio measures, where practical and statistical significance diverge, and why the computation should never come from the model being tested.

Read
Agents Multiply Reasoning. Ungoverned Agents Multiply Ungoverned Reasoning.
Agents Multiply Reasoning. Ungoverned Agents Multiply Ungoverned Reasoning.

A chatbot produces one reasoning event per interaction. An agent produces one per step — planning, tool selection, interpretation, revision — hundreds per task, invisible in the tool logs everyone mistakes for observability. The agent wave is a multiplication of exactly the thing enterprises never learned to govern.

Read