Updated: Colorado repealed SB 24-205 before it ever took effect and replaced it with the narrower SB 26-189, effective January 1, 2027. The original analysis stands on the record, with what the replacement dropped and what survives — including the question that outlives both statutes: can you document how the system reasoned?
Read →Research & announcements
Protocol releases, technical thinking, and updates from Arcus Labs
Almost there — check your inbox
We sent a confirmation link to finish your subscription. Click it and you’re all set.
Apply different prompt sets to the same reasoning graph and you get different convergence paths, different abstention rates, and different epistemic integrity scores. Prompt engineering isn’t cosmetic—it’s the configuration of an AI system’s epistemic personality.
Read →
Traditional validation assumes you can read a model’s mechanics, test it on holdout data, and stress-test it against known scenarios. LLMs break all three assumptions. What validation teams actually need isn’t holdout accuracy—it’s reasoning traces.
Read →
Hallucination rates aren’t evenly distributed. They increase with document complexity and input ambiguity—which means they concentrate in the domains where accuracy matters most: medicine, law, finance. The cause is architectural, and so is the fix.
Read →
Most EU AI Act coverage is written by lawyers for lawyers. But the Act’s requirements for high-risk AI systems aren’t just policy obligations—they’re architectural ones. Engineering teams need a different reading of Articles 9 through 15.
Read →
SR 11-7, the Federal Reserve's model risk management guidance, was written for statistical models with inspectable coefficients. LLMs break every assumption the framework rests on. When an examiner asks how the model arrived at a specific decision, the answer "we trust the output" is not an answer.
Read →