Request Demo

Research & announcements

Protocol releases, technical thinking, and updates from Arcus Labs

Occasional notes on new research and releases — a couple a month at most. We’ll send a confirmation link; unsubscribe anytime from any email with one click.

Almost there — check your inbox

We sent a confirmation link to finish your subscription. Click it and you’re all set.

Mapping a Reg E Dispute to a Reasoning Graph
Mapping a Reg E Dispute to a Reasoning Graph

The best-kept secret of graph design in regulated domains: the decomposition is already written. Regulation E’s error-resolution framework is a set of elements, each with its own evidence and its own clock — which is to say, it is a reasoning graph waiting to be transcribed. A worked example, element by element.

Read →
The First AI Was a Reasoner: Newell & Simon's Human Problem Solving
The First AI Was a Reasoner: Newell & Simon's Human Problem Solving

The founding programs of AI were engineered reasoners: explicit states, explicit operators, a difference list you could print and read. Newell and Simon’s General Problem Solver was running means-ends analysis by the late 1950s — every step inspectable by design. The statistical turn came decades later. In the history of artificial reasoning, next-token prediction is the newcomer, not the default.

Read →
The Transcript Is Not Evidence: What CoT Faithfulness Research Means for Examiners
The Transcript Is Not Evidence: What CoT Faithfulness Research Means for Examiners

Two research groups asked whether a model’s written reasoning actually explains its answer. Turpin et al. planted biases in prompts and watched answers shift while explanations never mentioned the bias; Lanham et al. cut and corrupted reasoning text and watched answers stay put. The lesson for anyone reviewing an AI decision file: a chain-of-thought transcript is testimony, not evidence.

Read →
Can AI Check Its Own Work? What the Self-Correction Research Shows
Can AI Check Its Own Work? What the Self-Correction Research Shows

A natural hope for language models is that they can catch their own mistakes: generate an answer, critique it, revise. A 2023 study from Google DeepMind tested that hope directly and found it wanting — without external feedback, self-correction did not improve reasoning and often made it worse. Reviewers of human work have known the underlying principle for a long time.

Read →
Reasoning Strategies, Part 3: The Scientific Family
Reasoning Strategies, Part 3: The Scientific Family

Problem-solving strategies find a path to a goal. The scientific family governs something harder: what you are entitled to believe, and how belief must move when evidence arrives. Two strategies, released as executable graph shapes: the hypothetico-deductive method and Bayesian updating.

Read →
The Compliance Officer's LLM Primer
The Compliance Officer's LLM Primer

You do not need the math. You need one accurate mental model: an assessor whose training is frozen into habits, whose entire knowledge of your case is the pile on the desk, and who answers by writing the most plausible next word. From that model, most of the governance questions answer themselves. This is lesson one of a curriculum.

Read →