In 2023, OpenAI researchers compared two ways of training a model to recognize good reasoning: reward the right answer, or reward each right step. Step-level supervision won decisively. Regulators will find the result unsurprising — it is how they have evaluated human decision-making all along.
Read →Research & announcements
Protocol releases, technical thinking, and updates from Arcus Labs
Almost there — check your inbox
We sent a confirmation link to finish your subscription. Click it and you’re all set.
The epistemic family asks what is true. The problem-solving family asks how to get from here to a defensible there — and its shared enemy is motion mistaken for progress. Five strategies, released together as executable graph shapes: decomposition, means-ends analysis, analogical reasoning, constraint satisfaction, and working backward.
Read →
The oldest information format still in production use is the Euclidean proof: definitions, assumptions, numbered propositions, each step warranted by something already established. We use it as an output contract for AI determinations — because it is the only prose format ever devised that makes missing justification structurally visible.
Read →
A natural experiment in regulatory architecture: take one model package, adjudicate it under SR 11-7 and under SR 26-2 on the same reasoning engine, and diff the dossiers. The evidence and math are identical; the classification, anchors, and conclusion vocabulary diverge. The diff is a map of where regulation actually lives.
Read →
AI-drafted SAR narratives are proliferating, and most of them optimize for exactly the wrong thing: fluency. A defensible narrative is not well-written — it is well-grounded: every factual claim traceable to the case file, the suspicion articulated as reasoning, nothing material omitted. Examiners are starting to ask how the narrative was produced. Institutions need an answer.
Read →
Effective challenge is the load-bearing idea in model risk management, and it has always run on organizational willpower — which is exactly what erodes under deadline pressure and familiarity. Encoding challenge as an adversary node makes it structural: steelman the position first, then attack in both directions, at a rigor locked to the model's materiality.
Read →