Someone hands you the reasoning record for a decision under review. Where do you look first? Not the beginning — read forward, a trace recruits you into the system’s own narrative. Here is a reading procedure in six passes, starting at the conclusion and working backward, and four things in a trace that should make a reader uneasy.
Read →Research & announcements
Protocol releases, technical thinking, and updates from Arcus Labs
Almost there — check your inbox
We sent a confirmation link to finish your subscription. Click it and you’re all set.
“AI-powered” describes the vendor’s architecture. It says nothing about yours — and your examiner will ask about yours. Here is the demand list that separates governable tools from black boxes: reasoning records as a contractual right, pinned versions, step-level evidence, provenance, sub-processors, examination support, and exit terms that leave you holding your own decisions.
Read →
In 1945, George Pólya wrote down what mathematicians actually do when they are stuck — and made the then-radical claim that it could be taught. Named heuristics, four phases, and a mandatory look back: the premise of a reasoning strategy library, stated eight decades early.
Read →
When an examination team asks an institution to reproduce an AI-assisted decision, what should count? Not “we asked the current system and got a similar answer” — that is an approximation, and it fails the standard. Replay means same inputs, same versions, same configuration, same record. Here is what must be pinned, and why a decision that cannot be re-run cannot be effectively challenged.
Read →
The first empirical failure taxonomy for multi-agent LLM systems reads less like a research artifact than an examiner’s checklist: 14 failure modes in 3 categories — bad specifications, agents talking past each other, and nobody checking the work — drawn from 150+ annotated execution traces. As agentic pilots arrive in banks, the taxonomy tells reviewers exactly where to look.
Read →