You Don’t Need Smarter Models
Inside many AI roadmaps sits an unstated plan: wait — the next model generation will be good enough to trust with regulated decisions. The plan confuses two axes. Capability — what the model can do — improves with scale. Defensibility — whether a decision can be traced, bounded, reproduced, and challenged — is a property of architecture, and scaling does not touch it. A more capable model in an ungoverned harness produces the same indefensible decision with better prose, which makes errors harder to catch, not rarer to matter. The institutions that build governed architecture now aren’t betting against model progress; they are building the only thing that makes model progress safely consumable.
The waiting strategy
Nobody writes it on a slide, but the reasoning is common enough to deserve a name. A team pilots an AI decisioning tool; the results are impressive but not trustworthy; the conclusion drawn is that the technology is almost there, so the prudent move is to wait a generation. Next year’s model will hallucinate less, reason better, and clear the bar. The waiting strategy treats trustworthiness as a capability threshold — a level the models will eventually reach, at which point the governance question answers itself.
The confusion is between axes. Ask what an examiner, an auditor, or opposing counsel will actually demand of a decision: show the path from evidence to conclusion; show what bounded it; run it again; let us challenge the step we dispute. Those are the four properties we’ve grouped before as defensibility — traceability, boundedness, reproducibility, challengeability — and none of them is a capability. They are properties of the machinery around the model: whether steps exist as inspectable units, whether anything blocks on failure, whether versions are pinned, whether a challenge has an artifact to operate on. A model cannot supply them at any scale, for the same reason a brilliant employee cannot supply your records-retention policy.
Eloquence raises the cost of doubt
Scaling does change something in the governance picture — for the worse. Detection of bad reasoning, by humans, runs on surface signals: hesitation, non-sequitur, visible confusion. Model progress systematically removes those signals while leaving some rate of substantive failure in place. The failures that remain arrive better-dressed — fluent, structured, confident — and the reviewer’s task shifts from noticing an error to auditing a persuasive account, which is slower, harder, and rarer. This is the pattern in the chain-of-thought faithfulness literature we read earlier this week: the explanation improves independently of the reasoning it claims to describe. An ungoverned harness converts model progress into persuasion progress, and persuasion is precisely the property a defensible process must refuse to rely on.
Waiting compounds the wrong debt
The waiting strategy also misprices time. Capability problems reward waiting: the next generation genuinely is better at the task. Governance problems compound against you: every ungoverned decision issued while waiting is a record that does not exist — a determination that cannot be reconstructed, a population that cannot be enumerated when a model change is questioned, an appeal with nothing to operate on. Institutions that waited out the last two model generations did not arrive anywhere; the bar they were waiting to clear is not on the capability axis. Meanwhile the examination environment moved — supervisory attention to AI-assisted decisions has only tightened — so the gap between what was deployed and what can be defended widened from both sides.
The inversion: governance is how you consume progress
Here is the part the waiting strategy gets exactly backward. A governed architecture — reasoning decomposed into scoped steps, gates that block, versions pinned per decision — is not a hedge against model progress. It is the adapter that lets an institution take model progress on board. When the reasoning is structured, a better model is a component swap: the graph, the gates, and the evaluation protocol stay; the substrate improves; the before/after comparison is measurable step by step, which is what change management requires anyway. We have made the deeper argument before — the architecture of the reasoning, not the substrate executing it, is what makes a decision inspectable — and it has a practical corollary: the governed institution adopts the smarter model in weeks, with evidence; the ungoverned one re-runs the same pilot and the same doubt.
Right-sizing follows from the same logic. Inside a graph, each step needs enough model for its job — retrieval steps, checking steps, and drafting steps have different requirements, and cheap generation under exact verification is often the better economics than maximum capability everywhere. The question stops being “is the frontier model smart enough to trust?” and becomes “which step still fails its gate too often?” — a question with a budget, a metric, and an owner.
None of this is an argument against better models; we will take every point of capability the labs ship. It is an argument about sequence. Capability is arriving on its own schedule, priced by someone else. Defensibility only arrives if you build it — and it is the piece that decides whether the capability is usable on decisions that answer to anyone.