Request Demo

Research & announcements

Protocol releases, technical thinking, and updates from Arcus Labs

Occasional notes on new research and releases — a couple a month at most. We’ll send a confirmation link; unsubscribe anytime from any email with one click.

Almost there — check your inbox

We sent a confirmation link to finish your subscription. Click it and you’re all set.

Grade the Steps, Not the Answer: What Process Supervision Shows
Grade the Steps, Not the Answer: What Process Supervision Shows

In 2023, OpenAI researchers compared two ways of training a model to recognize good reasoning: reward the right answer, or reward each right step. Step-level supervision won decisively. Regulators will find the result unsurprising — it is how they have evaluated human decision-making all along.

Read
Reasoning Strategies, Part 2: The Problem-Solving Family
Reasoning Strategies, Part 2: The Problem-Solving Family

The epistemic family asks what is true. The problem-solving family asks how to get from here to a defensible there — and its shared enemy is motion mistaken for progress. Five strategies, released together as executable graph shapes: decomposition, means-ends analysis, analogical reasoning, constraint satisfaction, and working backward.

Read
Euclid as an API: The Formal-Proof Layer
Euclid as an API: The Formal-Proof Layer

The oldest information format still in production use is the Euclidean proof: definitions, assumptions, numbered propositions, each step warranted by something already established. We use it as an output contract for AI determinations — because it is the only prose format ever devised that makes missing justification structurally visible.

Read
Same Model, Two Rulebooks: Adjudicating One Package Under SR 11-7 and SR 26-2
Same Model, Two Rulebooks: Adjudicating One Package Under SR 11-7 and SR 26-2

A natural experiment in regulatory architecture: take one model package, adjudicate it under SR 11-7 and under SR 26-2 on the same reasoning engine, and diff the dossiers. The evidence and math are identical; the classification, anchors, and conclusion vocabulary diverge. The diff is a map of where regulation actually lives.

Read
What Makes a SAR Narrative Defensible
What Makes a SAR Narrative Defensible

AI-drafted SAR narratives are proliferating, and most of them optimize for exactly the wrong thing: fluency. A defensible narrative is not well-written — it is well-grounded: every factual claim traceable to the case file, the suspicion articulated as reasoning, nothing material omitted. Examiners are starting to ask how the narrative was produced. Institutions need an answer.

Read
Effective Challenge as a Function
Effective Challenge as a Function

Effective challenge is the load-bearing idea in model risk management, and it has always run on organizational willpower — which is exactly what erodes under deadline pressure and familiarity. Encoding challenge as an adversary node makes it structural: steelman the position first, then attack in both directions, at a rigor locked to the model's materiality.

Read