Buried in the April guidance is the sentence almost nobody is writing about: generative and agentic AI models are out of scope. That is not an exemption. It leaves bank LLM deployments with no tailored framework, full supervisory exposure, and a question every examiner can still ask: show me how this system reached this decision.
Read →Research & announcements
Protocol releases, technical thinking, and updates from Arcus Labs
Almost there — check your inbox
We sent a confirmation link to finish your subscription. Click it and you’re all set.
The Digital Omnibus on AI cleared its final vote: the high-risk obligations that were due August 2 now arrive December 2, 2027. The content of the obligations survives; the deadline moved for standards readiness. The verified facts, what changed and what didn’t, and the implication that actually matters: the runway got longer, and undocumented decisions still don’t backfill.
Read →
The fastest way to run a big model turns out to be letting something small do most of the talking. Speculative decoding drafts tokens cheaply and verifies them in parallel — provably without changing the output. DeepSeek-V3 builds the drafter into the model itself and gets ~1.8× decoding speed. The pattern — cheap generation under strict verification — is one we recognize. The guarantee that makes it free at the token level is exactly what verification loses one level up, and that difference is worth understanding precisely.
Read →
Colorado repealed its landmark AI Act before it ever took effect and replaced it with something narrower: SB 26-189, built around automated decision-making technology, disclosure, and a consumer right to meaningful human review, effective January 1, 2027. The mandate list shrank. The question that survives — review of what, exactly? — is the one worth preparing for.
Read →
A reasoning model carries two kinds of momentum. One is worth keeping — the capability encoded in its weights by training. The other is the reason it can argue itself out of the right answer: the pull to stay consistent with whatever it said first. Chain-of-thought struggles to interrupt the second, because the check is written by the same running generation. A separate call can reduce it — and that, we argue, is most of why IRG is a different thing than a prompt.
Read →
VibeThinker-3B scores 94.3 on AIME26 with three billion parameters, matching models orders of magnitude larger. The claim underneath the benchmark is the interesting part: reasoning compresses aggressively while knowledge does not — meaning reasoning is a separable artifact. That has consequences for how production systems should be built.
Read →