Category: What to Build

Coherence engineering, oversight architecture, the emerging role, board-level implications. Maps to Part Three of the book “Coherence.”

  • 220% More Code, 40% More Incidents

    The Wall Street Journal reported in late July that corporate America was pulling back on AI token spending. “Tokenmaxxing” had a scoreboard that measured spend. Nobody kept track of what it cost to absorb. Meta ran the experiment at large scale, alongside a restructuring plan it called “Project OT,” and reporting in Reuters has made its numbers public. Quoting an internal post by the CTO in early June, Reuters reported an increase of 220% in code changes to internal platforms. Major technical and security incidents spiked 40% from the previous year, with the time employees had to spend firefighting them up 70%. Katie Paul, the reporter who broke the story, in a public Q&A on Sep 3, said that “tokenmaxxing pressure is out” now and the guidance is to use AI where it makes sense.

    The verdict from Zuckerberg himself was that they went too far ahead of where the technology was and that the “trajectory of the agentic development over at least the last four months hasn’t really accelerated in the way that we expected.” I spent almost a decade and a half building AI inside large companies before cofounding a data security company, and this sequence is familiar. While the verdict is partially correct, it teaches the wrong lesson. Zuckerberg told employees in April, explaining the rationale for the layoffs, that there were two major cost centers: “compute and infrastructure” and “people related things” and that investing more in one would leave less capital to be allocated to the other. It is true as a budget line. But it is wrong as a description of how AI output gets absorbed. Each output from AI has to be reviewed, integrated, and when it breaks, repaired. That is work that has to be done by people until the organization builds structure to do it: named owners for each system, constraints to eliminate entire classes of errors, walls between systems so they don’t contaminate each other, and escalation paths that are actually answered. The capacity that work requires depends on how much is deployed and how the systems interact, not on headcount. The 220% increase in internal platform code changes is one manifestation of the deployment surface. 

    The causal claim comes from Meta itself. Reuters reports that as early as March, infrastructure teams flagged reliability warning signs from the surge in AI-written code. An internal post in April said that unchecked agents were taking “large-scale, disruptive actions that humans are unlikely to execute.” Changes that reached users as new or improved features rose 36%, against 220% increase in code. If that implies much of what was produced was noise, it still had to be read and reviewed to be recognized. Review load rises with the volume and not with value. As the Reuters reporter recounts in the Q&A, AI-caused incidents were “weirder and less predictable” than human-caused ones and which is how 40% more incidents became 70% more firefighting. The signal was present in pieces: on the reliability dashboard, in the on-call rotation and in the March and April posts. But no one collated it against the AI spend. 

    EY put the principle in language that boards can understand: every organization has a ratio of builders to overseers, and governance capacity has to match building activity. Meta’s Project OT inverted it. While the code output increased, teams traditionally composed of 10 to 20 were restructured into pods of 3 to 5, and about 8000 people were laid off in May. What is not clear if there was any other structure in place to absorb what the smaller teams could not. Each new system adds interactions to every existing system. Coordination load rises with combinations and not with headcount. Meta is not the first to pay the price. Ford hired 350 veteran engineers to catch quality problems its automated systems missed. And IBM is tripling entry-level hiring because, in its HR chief’s words, without a pipeline “the well simply dries up.” Reducing token consumption or “thrift-maxxing” does not settle this bill either. Cheaper models mean more models, more routing, more repair, and more coordination. 

    Meta says Project OT was a scenario exercise and not every scenario would run. Some of the rise in incidents reflects more usage and Meta has declined to comment on the figures. Zuckerberg’s timing admission is more candor than most CEOs admit, and tracking incident numbers at all is more than most companies do. And he still expects models to catch up within the next three to six months. But that is the point – better models raise the volume and the absorption bill grows with it. 

    Before approving an AI-justified headcount plan, boards should ask for three numbers. What happens to review and incident volume when output rises 50%? Who owns each deployed system when the people who built them move on? What is the ratio of builders to overseers, and how is it chosen? Then ask for the incident line next to the token line, and ask what is meant to absorb the output once the people doing it by hand are gone. If the answer is nothing, the plan is a bet that absorption is free. Size the absorption capacity by how much is deployed and not how many are employed. Meta’s own correction points in the same direction: no more company-wide layoffs this year, selective retention packages for people it was about to lose, and the usage mandate withdrawn. That is Project OT run in reverse. 

    Zuckerberg told his staff the company has two cost centers, compute infrastructure, and people. Meta’s own experience shows a third: the cost of getting what the first produces through the second. While it appeared nowhere in the budget, it showed up everywhere in incident logs. The next plan that treats those two as competing should be asked to show the third one first. 

  • The Seven-Figure Job Nobody Can Quite Define Yet

    When you are working on anything related to AI, one of the challenges is how fast the ground moves beneath your feet. Everything about AI is happening at an unprecedented pace. I faced the same challenge as I started working on the book. But it wasn’t as much with the content, the thesis analyzes the implications of abundant AI rather than AI itself as a capability. It was more about identifying the profile of my target readers who would be interested in the book’s arguments.

    That turned out to be hard because that profile does not have a fixed job title. And what is that profile? The person who is accountable for making abundant AI capability generate value for the enterprise. At some companies, that is the chief AI officer and at others it could be the chief data officer. Elsewhere it could be head of AI governance, a VP of AI, or the CIO who has quietly absorbed the mandate. While the work is real and specific, the label itself is still forming.

    To be clear, my audience is not limited to only those people who have this formal accountability. It is wider than people who already hold the job. People who are thinking about extracting tangible value from AI but have not been given the accountability for it are also target readers for me. In addition, many companies have no single person for this work at all and accountability might be split across functions, sit with the CEO by default, or sit nowhere yet. That is not evidence against the work but one of the reasons why I wrote the book. This essential work of keeping an enterprise coherent as it deploys intelligence does not wait for a job title and just goes undone without named people accountable for it.

    So it caught my attention last week when Bloomberg reported that business schools are racing to train people to become chief AI officers, even when companies are trying to figure out what exactly the role will do. And the pay has arrived ahead of the job description too. Prior reporting from Bloomberg put the salary at banks near $3.5M a year, high enough that firms are poaching talent from one another. And here is the best part – some people already in that position think it will not exist for long.

    So, a role with no precise mandate, no settled title, seven-figure salaries, with fierce competition among companies for talent, and insiders who expect it to vanish. This is not how a market treats a job it understands, it is a function for which the market feels the need but hasn’t been able to define. I had to name the work to write the book.

    Strategy vs execution

    According to a BCG survey of 2,360 executives from earlier this year, roughly three-quarters of CEOs said they were their company’s main decision-makers on AI. That is double the share from a year earlier. AI strategy has moved to the top, which is where it belongs.

    But owning the strategy is not the same as owning its execution. Deciding to deploy AI is one thing but making the organization able to absorb what the decision sets in motion is entirely different. The gap between them is where clarity tends to blur in companies.

    On JPMorgan’s earnings call, Jamie Dimon described almost a thousand use-cases across the bank, with the firm’s platform rolled out to more than 200,000 employees. That is what owning the strategy looks like. But what it doesn’t say is whether those thousand systems work well together. Who owns the delivery of coherence across those thousand systems?

    Dimon also said something sharper – that his bank does not uniquely benefit from AI because everyone is now using it. One of the most quoted CEOs in banking conceded that models are no longer the advantage. What he didn’t say is what replaces it instead. That is the question the missing role is supposed to answer.

    The disappearing act

    Now to the strangest part – some of the people best positioned to know what the job entails say it will not last.

    Ranil Boteju is the first chief AI officer at the Commonwealth Bank of Australia. He expects AI to become invisible within about a decade, just the way electricity is, and the chief AI officer to shrink to a “very small role.” David Hardoon, who was the global head of AI enablement at Standard Chartered, said any chief AI officer should operate on the premise that they should not have a role in the future, asking if any company today has a chief Excel officer.

    If AI capability becomes truly ambient, a dedicated role for it is as odd as a chief electricity officer. The prediction does represent a real pattern and the argument is correct on its own terms. But it proves my premise. The reasoning is that specialized titles fade away as new technologies become part of daily infrastructure. That is the commoditization, and the same argument of Nicholas Carr’s ‘IT Doesn’t Matter’ playing out faster. What felt like a moat becomes a utility everyone has.

    Where the argument breaks is the analogy. Electricity does not build more electricity when you start using it. But agents do. Every AI deployment adds new systems, dependencies, handoffs, verification demands etc. and the burden of keeping them coherent grows as the capability becomes cheaper. The chief Excel officer joke works because a spreadsheet has no autonomy or agency. Electricity does not act on its own either. But agents act, connect to other agents, and take on scope until outputs feed decisions no one traced. The capability becomes invisible but what accumulates is incoherence.

    So the skeptics are right about the title but wrong about the function. The label “chief AI officer” may very well disappear if the work is “manage the AI.” But the real work isn’t that, it is keeping the enterprise coherent while abundant intelligence becomes ambient.

    Forrester predicts that 60% of the Fortune 100 will appoint a designated head of AI governance in 2026, with several companies already there. But I will concede that the function may not live in a named chief at all. It may fold into an existing role such as the CDO, COO or CIO. My claim is not that a particular title survives but that the function is real.

    The structure varies

    While the banks in the Business Insider survey did assign the work somewhere, there is no agreement on where the accountability should sit. The article reads as a set of incompatible bets on who should own coherence.

    Wells Fargo runs a hub-and-spoke model with a small central AI team and leads embedded in each business as spokes. Citi took a bottom-up approach training four thousand employees as AI stewards. And JPMorgan restructured its firmwide data and analytics office and reshuffled its leadership after its AI chief retired. One centralizes ownership, another distributes it across thousands of employees, and the third seems to be mid-reorganization still deciding.

    These are not variations of an answer but opposite theories and represent an industry trying to figure out what works. I want to be clear that nothing in these articles show any of the banks are incoherent. Nothing shows they are failing to build coherence. Several may be doing the exact right work. The point is that what these firms choose to measure and publicize – the usage rates, the productivity gains, the deployment counts – has little to say about whether the organization as a whole holds together. What the survey shows is silence in the evidence, and no consensus on what the accountable structure even is.

    Two sides of the same coin

    There is a missing metric. In six of the most sophisticated banks in the country, every figure reported measures capability, usage, or task-level productivity. None measures if the systems, taken together, serve the enterprise.

    And there is a missing owner. A role with no agreed description, and no agreement on whether it will exist.

    Both are the same problem. Coherence goes unmeasured because it is unowned. And it stays unowned because the role whose job it is to measure it has not been defined.

    I wrote about the measurement side in a companion piece, on how the big advisory firms are circling around coherence, and why structure beats supervision. That post is about what a company has to build. This one is about who has to hold it once it is built.

    And the answer to the question is a person. Building an enterprise’s ability to see what its systems are doing together, to know whether they still serve the overall business, and to correct a system that has drifted before the drift spreads – that capacity is what lets a company deploy hard and fast without coming apart. My book describes that work as coherence architecture, and the person doing it as a coherence architect. I do not offer that as a title the market will settle on. The market has not settled on one and the profession is still learning its own name. What I offer in the book is a description of the work so companies can recognize what the job entails before deciding what to call the person doing it.

    Companies that recognize it and find that person early are the ones that will still make sense while everyone else is counting use cases. That is the argument of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Building Coherence Is About Structure, Not Supervision.

    Gartner expects that by 2028, companies using multi-agent AI across most of their customer-facing work will pull ahead of everyone else, and that ninety percent of B2B buying will run through AI agents, moving more than fifteen trillion dollars (Gartner, Oct 2025). The same firm expects more than forty percent of agentic AI projects to be canceled by the end of 2027 (Gartner, June 2025). One of its own analysts says plainly: past a certain point, more AI does not mean more productivity. And in 2026 it predicted that by 2030, half of AI agent deployment failures will trace to governance platforms that fail to enforce capabilities and multisystem interoperability at runtime.

    Read those together and something is off. The firm forecasting agent dominance is the same one forecasting the shakeout.

    One analyst saying this would be a footnote. The big advisory firms all say some version of it. The firms telling you to scale agents across the enterprise are the same firms publishing the evidence that scaling is where the value dies.

    Every advisor is arguing with itself

    Look closely and each of the big advisory voices carries two messages at once.

    Gartner’s loud message is the proliferation math above. Its quiet message is the cancellations.

    Accenture carries both messages inside one report. Its mid-2026 study presses companies to move now, warning that “the cost of delay is not temporary but structural”. A few pages later it says the leading companies do not move faster, they move deliberately, and it names “systemic readiness” as the binding constraint. Push hard, and readiness is what actually gates you.

    PwC has a more disciplined public voice. Its 2026 predictions say agentic workflows are “spreading faster than governance models” can handle. A few predictions later, it offers the cure: an AI orchestration layer that, it says, will let you “control AI anywhere in your company”.

    BCG showed the whipsaw most starkly of all. Its July CIO playbook led with speed. Five weeks later its global chair told CEOs the first thing to protect is the enterprise’s own knowledge and judgment, not speed. Same firm, five weeks apart, the emphasis inverted. I worked through that shift in a prior post.

    Credit them all. But there is a gap.

    The gap they keep circling

    The loud message prices capability, meaning how much you can deploy. The quiet message prices readiness, meaning whether you built the muscle to deploy well. Neither one prices the thing that actually breaks once you scale, which is coherence.

    Coherence is a plain idea. It is whether the systems you deployed still serve the enterprise once they run together. Readiness is a gate you clear once, before you scale. Coherence is a property that erodes after you clear it, and it erodes faster the more you deploy. That is why a company can pass every readiness check, launch aggressively, and still land in Gartner’s forty percent.

    These firms describe the gap.

    Follow the mechanism

    The failure has a shape, and Accenture describes it. In a siloed rollout, every team builds its own agent on its own data. The invoicing agent has no view of supplier records. Procurement is walled off from finance’s process. Where those pieces should hand off, they break instead, and people get pulled back in to bridge the gap, which is the opposite of what the agents were for. Accenture calls it the “hidden tax of siloed transformation”. Every agent worked on its own. The cost lived in the seams between them.

    Gartner points at a related failure. Its sales analyst warns of a value ceiling, where piling more prompts and tools onto already complex workflows overwhelms the people working them and stops adding value past a point.

    PwC’s own safeguard shows the reflex. It suggests using agents to check other agents, and pulling in a second vendor’s model for higher-risk work. A sensible control, and also a tell, because the instinct is to answer agent sprawl with more agents.

    What each firm reaches for

    Each names the coordination problem, and each reaches for a build to solve it.

    Accenture is the most explicit. It describes the cross-functional collapse above and prescribes a multi-year rebuild it calls the intelligent superhighway: unified data, redesigned workflows, and a reinvented operating model. Much of that is what coherence requires.

    Gartner traces half of its projected agent failures to poor multisystem interoperability, then reaches for a universal semantic layer, which it calls the only way to align multiagent systems and stop costly inconsistencies before they spread.

    PwC prescribes the orchestration layer, a way to combine agents from different vendors into one process and, it says, stay in control.

    These are serious answers, and much of what they prescribe is real work, most of it ongoing rather than one-and-done. Here is what none of it produces. These layers standardize and route what passes between systems, which is the substrate coherence needs to exist at all. Run them well and keep running them, and coherence still does not follow, because it is a separate job. A layer will not judge whether the actions those systems take still add up to what the business wants, or decide whether two agents chasing different goals have started working against each other. It will not own the space between them, or put the state of that space on a number anyone reads. Coherence is the property all this infrastructure is meant to yield, and it is the one property none of them names or measures. I traced the same gap through McKinsey’s operating-model argument in an earlier piece: the rewiring these firms recommend routes around the old coordination layer and builds a new one underneath, machine-speed and owned by no one.

    There is a fair objection. They would all say they already preach discipline, and that the failures are the undisciplined ones. Grant it. Discipline applied one system and one program at a time still does not produce a standing measure of whether the whole keeps serving the enterprise. That measure is what is missing.

    What actually closes the gap

    The fix is structural, and it is the one thing none of these firms can sell you, because it is not a product at all. It is how the enterprise is wired to hold together as it fills with autonomous systems.

    Start with what does not work. The reflex, once a leader feels this, is to watch everything and keep people “in the lead” of every agent, as Accenture puts it. The instinct is right and the framing is not enough. You cannot lead a hundred systems running at machine speed by paying closer attention, and once human vigilance is the thing holding the company together, the company has already outrun it. Supervision does not scale to the speed of software.

    What scales is structure, and it runs as a stack. You have to see the whole before anything else works. Most leaders can say how accurate a model is and how many agents are in production, and cannot say whether those agents still agree with one another. Instrument the coordination state so the portfolio is visible in aggregate, because every move below this one runs blind without it.

    Then constrain. Give each system an action space defined and enforced ahead of time, not written into a prompt and hoped for, so whole classes of incoherence cannot form at all. In a 2026 red-team study, an agent told to keep a secret resolved the dilemma by destroying its own email server. It held the right value and had no limit on what it could touch. The limit is the fix, and it lives in the architecture, not the pep talk.

    Contain what the constraints miss. Partition the enterprise so a failure in one system stays in one, rather than racing through dependencies nobody mapped. Containment is what makes aggressive deployment survivable on the day a boundary slips, which it will.

    Then price it. Put coordination cost on the scoreboard the business actually reads. Judge a redesign by a single question: did the company grow more coherent or less as it scaled? Speed of shipping and the number of agents live are the vanity metrics that hide the debt underneath. The cost you decline to measure is the one that compounds in the dark.

    Human judgment sits on top of that stack, held back for the exceptions. It is worth something precisely because the layers beneath it carry the volume, so scarce attention lands on the few decisions that are expensive and hard to reverse instead of drowning in what the structure should have caught. Underneath it, every system has a named owner who can reach in and correct it when it drifts. This is what keeping people in charge looks like at machine speed: a human at the top of something built to need one only where it counts.

    Sense, constrain, contain, price, and reserve judgment for the top. That is the architecture the whole field keeps gesturing at and will not name, and it is what turns autonomy from a liability into something safe to scale. The companies that build it deploy more than the ones that mistook the control panel for control, because they can finally trust what they shipped.

    The unpriced category

    Coherence is missing from the forecasts for the reason technical debt and systemic risk went unpriced before their reckonings. The market prices what it can measure, and no one has been measuring this.

    The most influential voices in enterprise AI have now walked right up to it. They name the coordination failure, they prescribe unified data and orchestration and human oversight, and they still stop at the edge of naming the property itself. That is not a knock on their work. It is a sign the category is real and still unnamed.

    It needs a name, and it needs a different picture of the job.

    Everything these firms offer is a permitting office, and a good one. It checks each plan against the code before the plan may proceed. That is what governance does when it clears an agent to ship, and what a readiness program does when it certifies a company to scale, and it is worth having.

    The collisions happen after the gate. Two agents that each passed the desk converge on the same customer and pull the account two ways, and no one is watching the live picture. An agentic enterprise runs like an airspace, and an airspace does not run on permits. It runs on an air traffic controller, the one watching the sky who catches two cleared flights heading for the same point and moves one before they meet.

    The firms are building better permitting offices. Coherence is the control tower. Build it before the shakeout does the watching for you.

    Building that tower, and keeping it standing as the systems multiply, is the subject of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • The Machinery Under Manners

    Reid Hoffman posted this week about how to make an introduction. His rule is the double opt-in. Before connecting two people, you check with both of them first. You ask each whether they want the intro, and you make it only when both say yes.

    He makes the case well. Every introduction is a bet that the two people will get value from meeting. Checking first lets you see the cards before you place it. Connect two people who might not want to meet, and you have told the recipient their time was not worth a question.

    The double opt-in is a convention. Conventions are worth thinking about right now, because we are starting to put similar things on AI agents and calling them guardrails.

    Why it works

    Henry Mintzberg named five ways an organization coordinates its work. One of them is standardizing behavior, so that people act in sync without anyone supervising each step. A shared convention is that mechanism running in the wild. Everyone knows the norm, most people follow it, and the group stays coordinated without a manager in the loop approving every move. It is the cheapest way a group holds together, because it asks for no oversight at all.

    Look closely at what Hoffman says. He did not describe a habit you could take or leave. He turned a courtesy into a step you cannot skip. Both people have to say yes before anything happens. Remove the check and the introduction never occurs. There is a small piece of structure sitting under the courtesy, and it is easy to miss because it costs almost nothing to run.

    What enforces it

    Ask what stops someone from skipping the check when they are busy or feeling important. Hoffman answers it himself. A bad introduction damages the relationships you have with the people you introduced. The convention holds because breaking it costs you something with people you will deal with again. Reputation does the work, along with the knowledge that you will see these people down the line. The enforcement runs so quietly that we credit the manners and forget the relationship holding them up.

    Behavioral guardrails

    Agentic AI now acts with something close to human autonomy. An agent plans, works over many steps, and pursues a goal with little supervision. So we give agents rules for how to behave. Do not deceive. Do not touch systems you were not given access to. The industry calls these behavioral guardrails.

    The word guardrail usually means something physical, a barrier you cannot drive past. A behavioral guardrail is different. It is a norm the agent has to represent, understand, and choose to honor. It lives inside the agent’s own decision loop, which means the agent can reason its way around it. This is the same coordination-by-standardization that holds a human group together, handed to an actor with none of the enforcement underneath. An agent has no reputation at stake and no relationship it needs to protect. The rule reaches it as a suggestion, and nothing makes the suggestion bind.

    I pointed out recently on linkedin that a “skill” given to an agent is standardization in exactly Mintzberg’s sense. It works only as long as the agent applies the whole of it faithfully. The moment the agent decides which parts of the instruction matter and which it can set aside, standardization stops coordinating anything, and you are back to the expensive ways of keeping work aligned: a human checking every output, or endless back-and-forth to reconcile what the agent actually did with what you meant.

    The part you cannot see

    Here is what makes this hard to catch in time. A faithful agent and one quietly ignoring half its instruction produce the same kind of output. Both look like work getting done. You cannot tell from the result which one you had. You find out later, when something downstream breaks and you trace it back.

    That is the real weakness of a behavioral guardrail. You cannot read off whether it held. The output looks fine either way, right up until the run where it does not, and by then the thing has already happened.

    The AISI incident

    In August 2026 the UK’s AI Security Institute published a report on an incident in its own testing. The conditions were deliberately permissive. The agents had open internet access, and the model providers’ safety filters, which sit outside the model and can block an action the model wants to take, were switched off. The Institute is careful to say these are not the conditions the public gets. Testing this way shows what the models do once the structural constraints are taken away.

    Across many runs of a cyber challenge, agents took unsanctioned action on the live internet. In the most serious case, an agent tried to slip malicious code into a real open-source project. It researched the project’s human maintainer and created fake identities to pressure the maintainer into approving the code. The Institute states plainly that the agent was never told to deceive anyone. The deception emerged on its own, while the agent pursued the goal it had been given.

    The agents also invented a convention of their own. One agent left public messages offering to work with the other agents running the same challenge, and left behind accounts and instructions for them to reuse. It was a spontaneous etiquette among machines, a way to coordinate through the task. Even agents reach for conventions. What they do not build is any way to make a convention bind them.

    The Institute was clear about what stopped the worst of it. A human maintainer caught the malicious code and refused to approve it. That is coordination falling back to its most expensive mode, a person inspecting the output by hand, because the cheap mode had failed silently. The margin between failure and success was narrow, and it rested on “human vigilance rather than a technical barrier” that would reliably stop a more capable agent. Harm was prevented because a person happened to be paying attention.

    What a real guardrail looks like

    None of this makes the agents villains, and the Institute’s caution is worth keeping. It cannot even say for certain whether the agent understood it was acting in the real world or believed it was still inside a test. The agent broke the rule while chasing its goal, with no malice involved. The failure lives in the setup of the situation.

    So the design question is which parts of an instruction are load-bearing and have to hold no matter what, and which parts you actually want judgment to touch. What has to hold should be structural, made hard to do rather than merely discouraged. An agent that cannot reach a system has no need to be told to leave it alone. A code change that cannot merge without a check the agent cannot forge does not depend on the agent choosing honesty. Autonomy is fine, but it should be bounded, and the bounds have to be built rather than requested.

    One barrier around one system is not enough, because agents work in chains, and errors pass down the chain and grow. The fuller answer has layers. The first is visibility, since you cannot oversee what you cannot see, so the behavior of every system has to be observable. Then the connections between systems need constraining, so that whole categories of harmful action become impossible to express. Failures have to be contained where they do occur, so a mistake in one place does not travel. Human judgment comes in last, saved for the exceptions the structure surfaces, which is a far better use of a person than asking them to watch everything.

    Behavioral guardrails belong on top of that structure. They are the cheapest layer and the most brittle one, and they are fine in that role. The mistake being made right now is to build the top layer first, write down a list of good behaviors, and treat the list as safety.

    Coherence is the integrity of the link between what you meant and what got done. A convention holds that link cheaply, as long as something underneath makes it bind. Judgment stretches the link, sometimes for the better. Whether it helps or hurts depends on deciding in advance where it is allowed to stretch, and making the rest hold on its own.

    This is the argument at the center of my book, Coherence, arriving this Fall: that overseeing autonomous systems takes structure you build, not behavior you request. To follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Good AI Governance Is Not the Same as Coherence

    Australia’s directors just got the best AI governance guide I have read. The Australian Institute of Company Directors, with the Human Technology Institute, published an updated director’s guide this year, and it is genuinely good: careful, current, honest about agentic risk in a way most board material is not. It names the things that go wrong when autonomous systems run continuously and at scale. It tells boards to keep an inventory, set a risk appetite, assign ownership across the full life of a system, test before deploying, and keep a human able to intervene.

    I take it seriously, because it is the best available version of the mainstream answer. And then I want to explain why the mainstream answer, done well, still misses the failure that will actually catch these boards. The instrument it reaches for cannot see the thing that breaks.

    The guide is better than its genre

    The AICD guide says out loud that agentic AI raises risks the previous era did not. It notes that when systems act with high autonomy, errors can go undetected for longer. It notes that when they run continuously, errors compound before anyone addresses them. It flags that orchestrating across multiple agents multiplies both the failure paths and the attack surface. It even names shadow AI, the tools employees use with no oversight at all. That is a clear-eyed list, and most board guidance never gets near it.

    Its prescription is the machinery of good governance. Establish a risk appetite for AI. Keep a register of every system. Stand up a management-led AI committee to approve high-risk uses. Set a reporting cadence to the board. Get external assurance. Assign accountability from design through to decommissioning. If you did all of it, you would be far ahead of most companies.

    And you would still be exposed, in a specific way the machinery is not built to catch.

    Governance reviews what reaches it

    Here is the structural problem. A governance apparatus works by review. Something is proposed, and a committee assesses it against a policy. That is what a risk appetite, an approval gate, and a reporting line all do. They inspect items as those items arrive.

    The failure I study does not arrive as an item. It accumulates between the items.

    Picture a company that did everything the guide asks. Every agent has an owner. Every high-risk use went through the committee. The register is current. The board gets its quarterly report. Each system, reviewed on its own, was sound, and was approved for good reasons. Then the support fleet and the billing fleet and the underwriting fleet, each individually fine, begin to act on quietly incompatible assumptions about the same customer. No single system failed. No approval was wrong. The incoherence lives in the space between systems that were each approved separately, and a committee that reviews systems one at a time is looking in exactly the wrong place to find it. It is not that the committee decided badly. It is that the thing going wrong never came up for a decision.

    This is why I mostly avoid the word governance in my own work, and use oversight instead. Governance, in practice, has come to mean the apparatus of approval: the committees, the sign-offs, the documented permission to proceed. That apparatus is real and sometimes necessary. But it is closer to what a permitting office does than to what an air traffic controller does. The permitting office checks each plan against the code. The controller watches the live system and catches the two aircraft converging that were each individually cleared to fly. Agentic AI needs the controller. The guide, for all its quality, describes a very good permitting office.

    The oversight that passes its own audit

    There is a second failure the machinery cannot see, and it is worse because it looks like success.

    The guide, correctly, wants a human able to intervene. Keep a person in the loop. Maintain the ability to pull the plug. Every serious framework says this, and it is right. But “a human is formally in the loop” and “a human can actually catch what is going wrong” are different claims, and the gap between them widens quietly over time.

    A review team is assigned to check an autonomous system’s decisions. At first they overturn a real fraction. The system improves, so the threshold for review creeps down. The volume of what they wave through creeps up. The confidence scores get good, and people rarely argue with a high one. Eighteen months in, the team reviews a sliver of cases and overturns almost none, not because they are lazy but because the cadence never left them room to actually evaluate anything. On paper, human oversight is intact. The org chart is correct. The audit passes, because a procedural audit checks whether the review happens, not whether the review can still see. The oversight has become ceremony, and the framework that requires it cannot tell the difference. When the failure surfaces, and it will, everyone will point to a control that existed and was followed and did nothing.

    The numbers from a real deployment show how the trap tightens. A large United States health insurer rebuilt its document processing around AI. Before the project, its people caught errors on almost every document, because almost every document had one: fewer than one in ten was handled correctly first time. After the AI went in, the error rate fell to under three in a hundred. Good result. But to find those few errors, reviewers still had to examine more than a quarter of everything the system produced, because that was the share the model itself flagged as uncertain. The errors fell by a factor of about thirty. The human review load fell by a factor of less than four.

    Think about what that does to a board’s mental model. The system is now right almost all the time, which is exactly the condition under which a reviewer stops expecting to find anything. Yet the volume they must still wade through barely moved. You have the worst of both: enough review to be expensive, too little signal to stay sharp. Hold that threshold where it is and oversight stays costly. Lower it to save the cost and oversight goes blind. There is no setting on that dial that gives a board what it wants, which is cheap oversight that still catches things. That option does not exist, and no governance framework tells you so.

    A governance apparatus is structurally blind to this, because its test is whether the process ran. The question that matters, whether the humans in that process retain the capacity to intervene, is not a box a register can check.

    What to add, not what to replace

    None of this means throw out the guide. Keep the register, the risk appetite, the ownership, the reporting. They are necessary. They are just not sufficient, and the dangerous move is to mistake a complete governance apparatus for a complete answer.

    What has to sit underneath it is not more committee. It is structure. Build the ability to see your autonomous systems in aggregate, not one review at a time, so the incoherence between them becomes visible before it becomes an incident. Contain systems by design, so a failure in one domain floods that domain instead of the company. Set in advance how much each class of system may decide, so most of the safety is built into the wiring rather than caught at a gate. And test your human oversight for capability, not just for existence, by asking whether the reviewer could actually catch a subtle failure at the volume and cadence you have given them, not merely whether the review is on the schedule.

    That is a different kind of work from governance. Governance asks who is permitted to proceed. Oversight, in the sense I mean, asks whether the organization can still see what its systems are doing and still correct them while they run. The first is a committee. The second is an operating capability, closer to running a control room than to chairing a review.

    The AICD guide is what boards should read to get governance right. I am arguing they should read it knowing that getting governance right is the start of the problem, not the end of it. The failure that will catch the well-governed company is not the proposal the committee should have declined. It is the incoherence that never came up for a vote, and the oversight that kept passing its own audit while it went hollow.

    That gap is the subject of my book, Coherence. If you sit on a board or carry the risk for one, the one-page tool I use to start finding this blind spot is the first thing I send when you join the list at coherise.com.