Request Demo

Research & announcements

Protocol releases, technical thinking, and updates from Arcus Labs

Occasional notes on new research and releases — a couple a month at most. We’ll send a confirmation link; unsubscribe anytime from any email with one click.

Almost there — check your inbox

We sent a confirmation link to finish your subscription. Click it and you’re all set.

Pinned · Protocol Release Releasing irg-reference: An Open Implementation of the IRG Protocol
Can AI Check Its Own Work? What the Self-Correction Research Shows
Can AI Check Its Own Work? What the Self-Correction Research Shows

A natural hope for language models is that they can catch their own mistakes: generate an answer, critique it, revise. A 2023 study from Google DeepMind tested that hope directly and found it wanting — without external feedback, self-correction did not improve reasoning and often made it worse. Reviewers of human work have known the underlying principle for a long time.

Read
Reasoning Strategies, Part 3: The Scientific Family
Reasoning Strategies, Part 3: The Scientific Family

Problem-solving strategies find a path to a goal. The scientific family governs something harder: what you are entitled to believe, and how belief must move when evidence arrives. Two strategies, released as executable graph shapes: the hypothetico-deductive method and Bayesian updating.

Read
The Compliance Officer's LLM Primer
The Compliance Officer's LLM Primer

You do not need the math. You need one accurate mental model: an assessor whose training is frozen into habits, whose entire knowledge of your case is the pile on the desk, and who answers by writing the most plausible next word. From that model, most of the governance questions answer themselves. This is lesson one of a curriculum.

Read
Structure Over Surface: Gentner's Theory of Analogy
Structure Over Surface: Gentner's Theory of Analogy

In 1983, cognitive scientist Dedre Gentner formalized what separates a good analogy from a seductive one: good analogies map relations between things; bad ones match appearances. Every 'we had one just like this' in a case file is a bet on that distinction — and it is the exact distinction next-token prediction is worst at.

Read
What to Ask the Vendor: An Examiner's Question Bank
What to Ask the Vendor: An Examiner's Question Bank

Every vendor demo is fluent — fluency is the one property the technology guarantees. These twelve questions ignore the demo and probe the architecture: can the reasoning be inspected, re-run, bounded, and challenged? Each comes with what a good answer looks like, and what a bad one sounds like.

Read