Request Demo

Research & announcements

Protocol releases, technical thinking, and updates from Arcus Labs

Occasional notes on new research and releases — a couple a month at most. We’ll send a confirmation link; unsubscribe anytime from any email with one click.

Almost there — check your inbox

We sent a confirmation link to finish your subscription. Click it and you’re all set.

Measuring Reasoning Quality: Why Accuracy Benchmarks Miss the Point
Measuring Reasoning Quality: Why Accuracy Benchmarks Miss the Point

Accuracy benchmarks ask whether the system was right. That is the wrong question for production. The right question is whether the system behaved well given what it actually knew—and that is what epistemic integrity measures.

Read →
From the Perimeter to the Core: Why Auditability Unlocks AI Adoption
From the Perimeter to the Core: Why Auditability Unlocks AI Adoption

AI has already proven it can write, analyze, and classify. What it has not proven, in the institutional sense, is that its reasoning can be trusted for the decisions a regulated business is actually built on. The blocker is not capability. It is accountability.

Read →
The Colorado AI Act: What It Required — and What Replaced It
The Colorado AI Act: What It Required — and What Replaced It

Updated: Colorado repealed SB 24-205 before it ever took effect and replaced it with the narrower SB 26-189, effective January 1, 2027. The original analysis stands on the record, with what the replacement dropped and what survives — including the question that outlives both statutes: can you document how the system reasoned?

Read →
Prompt Sets as Epistemic Personalities: Same Graph, Different Reasoning
Prompt Sets as Epistemic Personalities: Same Graph, Different Reasoning

Apply different prompt sets to the same reasoning graph and you get different convergence paths, different abstention rates, and different epistemic integrity scores. Prompt engineering isn’t cosmetic—it’s the configuration of an AI system’s epistemic personality.

Read →
What Model Validation Looks Like When the Model Is an LLM
What Model Validation Looks Like When the Model Is an LLM

Traditional validation assumes you can read a model’s mechanics, test it on holdout data, and stress-test it against known scenarios. LLMs break all three assumptions. What validation teams actually need isn’t holdout accuracy—it’s reasoning traces.

Read →
Why AI Hallucination Rates Get Worse Where It Matters Most
Why AI Hallucination Rates Get Worse Where It Matters Most

Hallucination rates aren’t evenly distributed. They increase with document complexity and input ambiguity—which means they concentrate in the domains where accuracy matters most: medicine, law, finance. The cause is architectural, and so is the fix.

Read →