A natural hope for language models is that they can catch their own mistakes: generate an answer, critique it, revise. A 2023 study from Google DeepMind tested that hope directly and found it wanting — without external feedback, self-correction did not improve reasoning and often made it worse. Reviewers of human work have known the underlying principle for a long time.
Read →Research & announcements
Protocol releases, technical thinking, and updates from Arcus Labs
Almost there — check your inbox
We sent a confirmation link to finish your subscription. Click it and you’re all set.
Problem-solving strategies find a path to a goal. The scientific family governs something harder: what you are entitled to believe, and how belief must move when evidence arrives. Two strategies, released as executable graph shapes: the hypothetico-deductive method and Bayesian updating.
Read →
You do not need the math. You need one accurate mental model: an assessor whose training is frozen into habits, whose entire knowledge of your case is the pile on the desk, and who answers by writing the most plausible next word. From that model, most of the governance questions answer themselves. This is lesson one of a curriculum.
Read →
In 1983, cognitive scientist Dedre Gentner formalized what separates a good analogy from a seductive one: good analogies map relations between things; bad ones match appearances. Every 'we had one just like this' in a case file is a bet on that distinction — and it is the exact distinction next-token prediction is worst at.
Read →
Every vendor demo is fluent — fluency is the one property the technology guarantees. These twelve questions ignore the demo and probe the architecture: can the reasoning be inspected, re-run, bounded, and challenged? Each comes with what a good answer looks like, and what a bad one sounds like.
Read →