Australia’s directors just got the best AI governance guide I have read. The Australian Institute of Company Directors, with the Human Technology Institute, published an updated director’s guide this year, and it is genuinely good: careful, current, honest about agentic risk in a way most board material is not. It names the things that go wrong when autonomous systems run continuously and at scale. It tells boards to keep an inventory, set a risk appetite, assign ownership across the full life of a system, test before deploying, and keep a human able to intervene.
I take it seriously, because it is the best available version of the mainstream answer. And then I want to explain why the mainstream answer, done well, still misses the failure that will actually catch these boards. The instrument it reaches for cannot see the thing that breaks.
The guide is better than its genre
The AICD guide says out loud that agentic AI raises risks the previous era did not. It notes that when systems act with high autonomy, errors can go undetected for longer. It notes that when they run continuously, errors compound before anyone addresses them. It flags that orchestrating across multiple agents multiplies both the failure paths and the attack surface. It even names shadow AI, the tools employees use with no oversight at all. That is a clear-eyed list, and most board guidance never gets near it.
Its prescription is the machinery of good governance. Establish a risk appetite for AI. Keep a register of every system. Stand up a management-led AI committee to approve high-risk uses. Set a reporting cadence to the board. Get external assurance. Assign accountability from design through to decommissioning. If you did all of it, you would be far ahead of most companies.
And you would still be exposed, in a specific way the machinery is not built to catch.
Governance reviews what reaches it
Here is the structural problem. A governance apparatus works by review. Something is proposed, and a committee assesses it against a policy. That is what a risk appetite, an approval gate, and a reporting line all do. They inspect items as those items arrive.
The failure I study does not arrive as an item. It accumulates between the items.
Picture a company that did everything the guide asks. Every agent has an owner. Every high-risk use went through the committee. The register is current. The board gets its quarterly report. Each system, reviewed on its own, was sound, and was approved for good reasons. Then the support fleet and the billing fleet and the underwriting fleet, each individually fine, begin to act on quietly incompatible assumptions about the same customer. No single system failed. No approval was wrong. The incoherence lives in the space between systems that were each approved separately, and a committee that reviews systems one at a time is looking in exactly the wrong place to find it. It is not that the committee decided badly. It is that the thing going wrong never came up for a decision.
This is why I mostly avoid the word governance in my own work, and use oversight instead. Governance, in practice, has come to mean the apparatus of approval: the committees, the sign-offs, the documented permission to proceed. That apparatus is real and sometimes necessary. But it is closer to what a permitting office does than to what an air traffic controller does. The permitting office checks each plan against the code. The controller watches the live system and catches the two aircraft converging that were each individually cleared to fly. Agentic AI needs the controller. The guide, for all its quality, describes a very good permitting office.
The oversight that passes its own audit
There is a second failure the machinery cannot see, and it is worse because it looks like success.
The guide, correctly, wants a human able to intervene. Keep a person in the loop. Maintain the ability to pull the plug. Every serious framework says this, and it is right. But “a human is formally in the loop” and “a human can actually catch what is going wrong” are different claims, and the gap between them widens quietly over time.
A review team is assigned to check an autonomous system’s decisions. At first they overturn a real fraction. The system improves, so the threshold for review creeps down. The volume of what they wave through creeps up. The confidence scores get good, and people rarely argue with a high one. Eighteen months in, the team reviews a sliver of cases and overturns almost none, not because they are lazy but because the cadence never left them room to actually evaluate anything. On paper, human oversight is intact. The org chart is correct. The audit passes, because a procedural audit checks whether the review happens, not whether the review can still see. The oversight has become ceremony, and the framework that requires it cannot tell the difference. When the failure surfaces, and it will, everyone will point to a control that existed and was followed and did nothing.
The numbers from a real deployment show how the trap tightens. A large United States health insurer rebuilt its document processing around AI. Before the project, its people caught errors on almost every document, because almost every document had one: fewer than one in ten was handled correctly first time. After the AI went in, the error rate fell to under three in a hundred. Good result. But to find those few errors, reviewers still had to examine more than a quarter of everything the system produced, because that was the share the model itself flagged as uncertain. The errors fell by a factor of about thirty. The human review load fell by a factor of less than four.
Think about what that does to a board’s mental model. The system is now right almost all the time, which is exactly the condition under which a reviewer stops expecting to find anything. Yet the volume they must still wade through barely moved. You have the worst of both: enough review to be expensive, too little signal to stay sharp. Hold that threshold where it is and oversight stays costly. Lower it to save the cost and oversight goes blind. There is no setting on that dial that gives a board what it wants, which is cheap oversight that still catches things. That option does not exist, and no governance framework tells you so.
A governance apparatus is structurally blind to this, because its test is whether the process ran. The question that matters, whether the humans in that process retain the capacity to intervene, is not a box a register can check.
What to add, not what to replace
None of this means throw out the guide. Keep the register, the risk appetite, the ownership, the reporting. They are necessary. They are just not sufficient, and the dangerous move is to mistake a complete governance apparatus for a complete answer.
What has to sit underneath it is not more committee. It is structure. Build the ability to see your autonomous systems in aggregate, not one review at a time, so the incoherence between them becomes visible before it becomes an incident. Contain systems by design, so a failure in one domain floods that domain instead of the company. Set in advance how much each class of system may decide, so most of the safety is built into the wiring rather than caught at a gate. And test your human oversight for capability, not just for existence, by asking whether the reviewer could actually catch a subtle failure at the volume and cadence you have given them, not merely whether the review is on the schedule.
That is a different kind of work from governance. Governance asks who is permitted to proceed. Oversight, in the sense I mean, asks whether the organization can still see what its systems are doing and still correct them while they run. The first is a committee. The second is an operating capability, closer to running a control room than to chairing a review.
The AICD guide is what boards should read to get governance right. I am arguing they should read it knowing that getting governance right is the start of the problem, not the end of it. The failure that will catch the well-governed company is not the proposal the committee should have declined. It is the incoherence that never came up for a vote, and the oversight that kept passing its own audit while it went hollow.
That gap is the subject of my book, Coherence. If you sit on a board or carry the risk for one, the one-page tool I use to start finding this blind spot is the first thing I send when you join the list at coherise.com.
Comments
One response to “Good AI Governance Is Not the Same as Coherence”
[…] Good AI Governance Is Not the Same as Coherence. Australia’s directors just got one of the best AI governance guides I’ve read, and I spent the post explaining why the best version of the mainstream answer still misses the failure that will catch these boards. A governance apparatus works by review, and it reviews what reaches it. The incoherence that builds up between separately approved systems never comes up for a vote. And the human oversight everyone prescribes can pass its own audit while quietly going hollow, as a stretched review team waves through more and catches less. Real news from last week makes the point. Anthropic, the company that sells agentic AI, published a sober four-question checklist for deploying it safely: what untrusted content does the agent ingest, what can it do, what’s the blast radius, can you see what it’s doing. Four good questions. Every one of them inspects a single agent, one at a time. None of them can see the incoherence that accumulates in the space between agents that each passed. [Weigh in on LinkedIn…] […]