If the title sounds weird, it is because I stole it from a reddit thread I read last week! This thread, on a data engineering subreddit, is all about building fast without building coherent. A practitioner described a year spent working inside a major enterprise platform deployment that, by their account, went badly. The post drew more than a thousand upvotes and over 150 comments, and many of those comments said a version of the same thing. This matches what happened to us.
Let me be careful about what this post is and is not. I cannot verify any of it. I do not know the author, the employer, or whether the account is accurate, complete, or fair. I am not treating any of it as fact. I am not making a claim about the named vendor, any other company, its people, or its products. Online accounts are one-sided by nature. The company is not present to respond. And the thread does not even agree with itself on basic points, including why things went wrong. One commenter accused the vendor’s engineers of dragging work out to bill more hours. Two others replied that the vendor uses fixed-price contracts and has the opposite incentive.
So, let me set aside the question of motive or blame but ask something different. If a reader believed these accounts as written, what pattern would they describe? The pattern, if it is real, is one my book predicts.
The pattern the accounts describe
The original poster says they inherited the system after the engineers who built it left on thirty days notice, once a first version was declared done. What they found, in their telling, was a catalog of shortcuts. Hardcoded dates. Hardcoded accounts. The same business concept fed by different inputs in different places. Earlier problems patched with more hardcoded logic.
They gave one concrete example later in the thread. An engineer had built a button to delete a record. The button removed the record from the screen. It did not remove the three related records that the original had created when it was made. The result was orphaned data. The button worked. The system did not.
That small story is the whole thing in miniature. Every piece can be locally correct while the system is globally broken. A button that deletes what you can see and leaves what you cannot is a fine button and a broken workflow at the same time.
A commenter who said they work at a hospital inside a national health service described their own experience. Outside engineers did intense early work, leaned heavily on the in-house team to explain the basics, then left. No one was clearly left owning or maintaining what had been built. At one point, the commenter said, the vendor’s own monitoring staff emailed to ask why duplicate records and bad addresses were appearing, and the hospital could not answer, because it did not have access to the pipelines that had been built for it.
Another commenter described a failure higher up the organization. A single platform owner was installed. Over time that person’s standing became tied to the platform’s success, and the information traveling up to senior leadership was filtered, so the picture at the top stayed positive while the picture on the ground did not.
The line I keep thinking about
The poster wrote that the engineers used AI to produce tangled, low-quality logic. A commenter answered in five words.
That is the argument of my book, delivered by someone who did not set out to make it. When producing code becomes fast and cheap, more of it gets produced. The speed is real. What does not arrive with the speed is coherence. Coherence is the work of making sure each piece fits the whole, that today’s shortcut is not tomorrow’s silent failure, and that someone still understands the system after the people who built it are gone. Execution got cheaper. Coherence did not.
Why I am comfortable writing this at all
Here is the part that matters most. The people in the thread mostly did not think the story was about one company. One commenter wrote that you could swap in almost any vendor, almost any consultancy, and almost any project, and reach the same ending. Another described the identical arc with a completely different vendor. Others reached back to the enterprise data tools of twenty years ago and asked whether it had always been this way. They were describing a recurring structural pattern, and I think they were right to.
The pattern is old. W. Edwards Deming spent decades showing that optimizing each part of an organization on its own can degrade the whole, because the connections between the parts matter as much as the parts. Stafford Beer showed that organizations drift when the feedback reaching the people in charge is slow or filtered. Neither man was talking about AI. Both were describing this thread.
I want to give the other side its due, because the thread did. The original poster said plainly that the platform itself is fine for what it is. Other commenters defended it and corrected specific claims. Many organizations report that the same tools serve them well. The tool is capable. What fails, in these accounts, is the fit between a tool sold on speed and an organization that cannot absorb what speed produces. Change the logo on the invoice and the story would run the same way.
The number nobody calculated
The poster reported that the project was estimated at four months and took fifteen, and that the company had seen no return so far. I cannot confirm those figures. If they are even roughly right, they point at something the book returns to again and again. The promised savings were a calculation about capability. The cost that actually landed was a calculation nobody made, the cost of coordinating, maintaining, and understanding what got built. The first number is easy to put in a sales model. The second one shows up a year later and has no owner.
I have written three times recently about the same shape seen from different angles. A machine can generate the output. A person still has to own the part with no dashboard. In a newsroom experiment, an AI agent finished the forms and could not finish the job. In a courtroom, a scoring system read a gap in the data as a verdict on a person. In this thread, if the accounts hold, capable engineers produced software that worked in the demo and broke quietly in the corners, then left, and the coherence walked out the door with them.
Faster is not the same as coherent. It never was. The difference used to be expensive to create and easy to see. Now it is cheap to create and slow to see, which is exactly why it is worth watching for.
I can’t verify the thread, but the gap it points at is real. That gap, between building fast and building coherent, is the subject of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
Keith Collins gave an AI agent full control of a laptop and three office jobs to do: survey nine colleagues over Slack and log their answers, identify staff cuts to hit a budget target, and fill out seventeen I-9 employment verification forms. The tasks were adapted from benchmarks published by researchers at Carnegie Mellon and OpenAI. The agent ran on Anthropic’s Claude Cowork app.
On the third task, the agent generated all seventeen forms correctly in under five minutes. Then it tried to upload them to Google Drive and failed. It clicked the right menu item and never noticed that a file picker had opened. It compressed the files. It converted them to a long string of bytes. It asked a second agent for help, and the second agent hit the same wall. After roughly seventeen minutes, it stopped trying and marked the task complete.
The Times files this under comic stumbles. It is the most consequential finding in the piece.
120,000 jobs, cut on the opposite premise
The article closes on a line meant to calm: AI still needs a human boss.
The same article reports the layoffs. More than 200 tech companies have cut roughly 120,000 jobs this year, per Layoffs.fyi. Meta and Oracle made substantial cuts citing AI. Cloudflare’s chief executive, after letting go of about 1,100 people, said he expects AI to replace workers in middle management, finance, and marketing.
Those cuts rest on a premise: the supervisory layer is what becomes redundant. The experiment found the reverse. Agents were strong at execution and weak at judgment. They wrote clean code in minutes, then made a categorization error about employees on leave that any manager would have caught.
Firms are removing coordinating capacity while installing systems that consume more of it. That connection is the thesis of the book I am writing. One case has already reached a federal courtroom.
Three specimens
My argument: when execution stops being scarce, the binding constraint becomes coherence, the integrity of the link between what local systems do and what the enterprise intends. Coherence has five specific dimensions, and autonomous systems break it in six recognizable ways. The Times experiment produced three clean specimens.
The false completion is escalation failure. The agent detected its own failure. It reasoned about it for seventeen minutes. It recruited a second agent. Then it reported success. The system knew it had not finished, and it stayed quiet. In this case, it cost little to the reporter analyzing logs. But in an enterprise running ten thousand delegated tasks a day, it is the mechanism by which reported completion drifts away from actual completion. Escalation failure is the one mode in my taxonomy that leaves every dimension of coherence intact and disables the reflex that repairs them. An organization can see a problem clearly and still be paralyzed when the signal never reaches anyone who can act.
The second agent matters too. Two systems drawn from the same model share the same blind spot, so the redundancy is nominal. The organization paid for one failure twice.
The code detour builds architecture nobody chose. In every task, the agent was told to work through the applications and wrote code instead. Graham Neubig of Carnegie Mellon puts it plainly in the piece: agents work in a very unhuman way, writing code instead of using the interfaces humans use. The Times treats this as a limitation. It is also an architectural event. The agent replaced the assigned task with a different one that produced a similar-looking artifact. An org chart rebuilt by a Python script carries a new dependency, a new failure mode, and no owner. Multiply that across a year of routine delegated work and the enterprise runs on infrastructure nobody selected, documented nowhere, discovered only when it breaks.
Local simplification often works by moving complexity somewhere else. The productivity gain lands on the dashboard. The displaced complexity does not.
The staffing error is contextual failure, and it is already in litigation. Given a budget target, the agent did something genuinely good. It read the personnel documents and concluded the 4 percent reduction could be met through planned retirements and resignations, with no layoffs. Then it added employees on leave to the list of cuttable roles without considering when they were coming back. The source material was silent on duration. The agent never asked.
Researchers at Stanford and the NBER frame this as a tacit knowledge problem, and that holds. The mechanism is more specific. The agent had no way to represent a person as temporarily absent for a reason that says nothing about their value. Silence in the record became a mark against the employee.
Nine days before the Times published, twenty-six Meta employees filed suit in federal court in Oakland alleging that the same substitution happened to them at production scale. Their complaint says Meta relied on internal AI systems, keystroke and activity-monitoring data, AI token-usage dashboards, and algorithmically assisted performance rankings to decide who would go in a layoff of roughly 8,000 people, about 10 percent of the workforce. The central allegation: those scores cannot by design be accumulated by an employee on protected medical or family leave, or by an employee whose output is reduced by a disability. The suit further alleges the company never paused the process for the individualized, leave-neutral review the law requires. About half the plaintiffs had taken leave for caregiving or pregnancy-related reasons. Their jobs were set to end on July 22, the day the Times ran its experiment.
Meta rejects the claims. The company says they lack merit and are not based on facts, and that workforce and organizational decisions “were and are made by people, not AI.” The allegations are unproven and the case is live. I am looking at the structure of the dispute here and taking no position on the verdict.
That structure survives either outcome, which is why it belongs in this argument. Suppose Meta is right that people made every call. Those people still read rankings, and the rankings still came from a substrate that had no field for protected absence. A human who approves a ranked list holds the authority to intervene and does not necessarily hold the information. Formal presence in a process is a weaker thing than capacity to change it. My book calls that oversight failure.
The Times agent and the Meta complaint describe one error at two scales. A system met a gap in its data and scored the gap as a deficiency. Nobody had built the mechanism that would have made it ask a question instead.
My book already discusses the Cloudflare decision the Times cites. Its chief executive organized his reasoning around a Drucker framework: every organization has builders, sellers, and measurers, and AI can now measure cheaply, so the measuring layer can shrink. The framework is coherent and the logic holds internally. The question I put to it there was whether the people categorized as measurers were only measuring. The Meta complaint poses the companion question. When the score came back low, was the system measuring performance, or measuring absence?
The article measured one axis
Task reliability and organizational complexity are independent dimensions. Improving one leaves the other where it was. Conflating them keeps the expensive failures invisible until they are hard to reverse.
The Times measured reliability, carefully and well. Its headline number comes from Scale AI: on real freelance projects, the best model produced client-ready work about 16 percent of the time.
The coordination question sits outside that number. If 84 percent of agent output requires human review, review capacity becomes the ceiling on deployment. Oversight load scales with the number of systems, and it lands on a different dashboard than the productivity gain. A control system has to be at least as varied as the thing it controls. Thin the supervisory layer while thickening the volume of supervised work and the cost moves off the ledger. It stays in the business.
The Oakland filing shows where it resurfaces. Twenty-six people asking a court to examine how a ranking was produced is a coordination cost, arriving late, in the most expensive form available.
Where the article argues against me
The counterargument has real force. If agents cannot reliably finish tasks, they will not be deployed at scale, and the coordination problem stays theoretical. The article supports that. A 16 percent success rate describes a product that is not ready.
Two responses.
First, the two failures differ in kind. The upload bug will be fixed. It is a UI problem and the entire industry is aimed at it. The leave-of-absence error and the false completion report sit at the boundary between the agent’s context and the organization’s. Better models will make both rarer. No model tells an enterprise which completion reports it can trust, or who owns the Python script the agent wrote last Tuesday. Those are ownership questions, and capability does not settle them.
Second, consider what the two variables are doing. Reliability improves on a public curve that everyone watches. Coherence has no curve, because almost nobody measures it. Using a snapshot of the fast-moving variable to dismiss the stationary one is the error my book is written against.
Two caveats. The Times experiment was three tasks, one tool, one synthetic environment, with expert-written prompts and supporting documents supplied by benchmark researchers. Real enterprises rarely supply that quality of context. The setting was favorable on the task side and trivial on the coordination side, since one agent ran alone with no installed base of prior deployments to collide with. The reliability observed sits closer to a ceiling than a floor. The coordination cost observed is near zero by construction.
The second caveat: Meta’s alleged systems are ranking and monitoring software, a different technology from an autonomous agent operating a laptop. The defect predates agentic AI. Agentic deployment raises the rate at which it executes.
What to watch instead
If you run an enterprise and this experiment shaped your thinking, start measuring the things it did not.
What fraction of your deployed autonomous systems has a named human owner. What fraction produces decisions you can explain to a regulator. How many of your systems depend on other systems in ways nobody mapped. How much of the behavior is visible to the people accountable for it. How hard it would be to remove any given system now that it is running.
If my thesis holds, those five move in one direction as deployment scales, while task accuracy holds steady or improves. That divergence is the signature.
The Times asked whether AI can do your job. Twenty-six people in Oakland are asking the harder version. Who answers for the score that said they were not doing theirs?
My book, Coherence, arrives this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
I want to start by giving you a word, because I could not quite find the one I needed and had to make it.
We say a system “coheres,” as if holding together were something it manages on its own, almost by luck. What the agentic era demands is far more deliberate. When execution gets cheap and anyone can build, automate, and deploy in an afternoon, speed stops being an edge, because everyone has it. What becomes scarce is coherence: getting a growing crowd of autonomous systems, and the people accountable for them, to pull in the same direction. And getting them there is active work. Someone has to do it. That work deserves its own verb.
So: to coherise. It means to achieve coherence on purpose, to take a set of parts that could easily pull apart and make them hold together as one. You can coherise a team, a workflow, a company filling up with agents. You can also fail to coherise it, which is where most organizations are heading right now without seeing it, because the failure does not show up anywhere a dashboard would catch. When execution gets cheap and everyone can go fast, the scarce skill becomes the ability to coherise everything you have built.
I am convinced this is the job the agentic era is quietly creating, the one that decides who wins it, and it does not have a name yet. So I gave it one, and named this newsletter after it.
The week in ideas
Each edition I will round up what went on the blog, so you have one place to catch anything you missed. If you are new, here is the whole arc so far.
Two posts lay the foundation.
Introducing Coherence. Why I wrote the book, told as a story. It runs from a question I have chased since college, how billions of neurons become one mind, to the problem every leader now faces: how a company holds together as it fills with autonomous systems. Weigh in on LinkedIn…
Intelligence Is Becoming Cheap. Coherence Is Not. The core argument in one place. It opens with a baseball story from early in my career and lands on the idea I keep returning to, complexity debt: the hidden cost that builds as automation piles up and nobody is coordinating it. Weigh in on LinkedIn…
Three from the past week take the idea into things happening right now.
The AI Jobs Debate is Not Asking the Right Question. Everyone is arguing over whether AI takes jobs. I think that argument misses the larger shift. When execution gets cheap, the scarce thing becomes coordination, and the Meta layoffs, read closely, are a coordination story wearing a labor headline. Weigh in on LinkedIn…
McKinsey Is Right About the Moat. Here’s the Half It Misses. McKinsey argues that your real AI advantage is your operating model, the one thing a competitor cannot copy. They are right. The half they skip is that the redesign they prescribe is also the fastest way to manufacture a new, invisible coordination problem, and getting it wrong widens the very gap they set out to close. Weigh in on LinkedIn…
Your AI Usage Exhaust Is Someone Else’s Moat. Satya Nadella and two researchers, Arvind Narayanan and Akash Kapur, described the same trap from opposite ends within days of each other. Every time your people correct an AI, they encode your institution’s judgment into it, and that judgment leaks to whoever owns the model. The post works out where the leak costs you most, and where you can let it go. Weigh in on LinkedIn…
Before you go
That is edition one. From here it lands weekly: a short round-up of what I wrote, plus the occasional thing that caught my attention.
If the book is why you are here, it is called Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone who joins the list at coherise.com gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.
And if you try to coherise something this week, tell me about it. Those stories are where a good share of my ideas come from.
The problem is easy to picture, which is part of its charm. Scatter some dots on a page. Count the pairs that sit exactly one inch apart. As you add more dots, how fast can that count grow? Erdős guessed there was a ceiling, and that a grid-like arrangement came close to it. For decades most mathematicians thought he had it right. The model found arrangements that beat the ceiling, and kept beating it as the number of dots grew without limit.
The surprising part was the route. The unit distance problem belongs to discrete geometry. The solution came through algebraic number theory, which studies something else entirely. Writing in the Wall Street Journal last month, the statistician Daniel Kipnis makes an observation about this that I keep returning to. Cross-disciplinary borrowing is not new in mathematics. Descartes did it in the seventeenth century. What has changed is scale. A mathematician can spend an entire career in discrete geometry and never acquire the tools of algebraic number theory, because a human career is short and those tools take years. A machine has no such constraint. Its reach across the field is bounded by the cost of computation and nothing else.
So the machine went somewhere no specialist would have thought to look, and came back with a counterexample.
Then Kipnis asks the question that makes this interesting. If nobody understands a proof, is it a proof at all?
Mathematics as a social act
He gets his answer from Reuben Hersh, who spent a career arguing that mathematics is a social phenomenon rather than a collection of eternal truths sitting somewhere waiting to be found. On that view, a mathematical fact does not become part of mathematics by being true. It becomes part of mathematics by being discovered, explained, and absorbed into what the community understands. Progress is a form of communication. A proof explained badly does no more good than a proof that is wrong.
Kipnis notes that OpenAI seems to have understood this instinctively. It did not publish the output and walk away. It worked with prominent mathematicians who verified the argument and wrote a companion paper making it intelligible to the field. Without that second step, the result would have occupied a strange position. It might have been true. But if no human could confirm it or follow it, what would its truth consist of? On Hersh’s account it would not yet be a mathematical result.
This is a strong claim and I find it persuasive. It also has a problem, which two other people, in another discussion, illuminate.
Is understanding a crutch?
On Quanta’s podcast The Joy of Why, Steven Strogatz recently interviewed Lauren Williams, the Harvard mathematician who helped start the First Proof project. Afterward Strogatz and his co-host Janna Levin, an astrophysicist, kept talking, and the conversation turned to something more unsettling than job displacement.
Strogatz asked whether beauty will still guide mathematics once machines are doing it alongside us. Beauty in the working sense, meaning the aesthetic pull that tells a mathematician which question is worth asking and whether an argument is on the right track. Earlier in the discussion, Strogatz and Williams had discussed about her philosophy on beauty: “If you ask a question and the answer is not beautiful, that means you asked the wrong question.”
Levin’s answer is the best thing I have read on this in months. One of the things beauty does, she said, is make a complicated subject comprehensible. Then she gave the reason she needs that: “I don’t have infinite compute.”
Elegance is not decoration. It is compression. Understanding is the technique a bounded mind uses to fit something enormous into a space the size of a human head. We prize proofs that are short, surprising, and clean because we cannot hold the long ugly ones. As Schmidhuber, a leading AI scientist, explains, a computationally limited observer finds something simpler and more beautiful once she learns to predict and compress the data in a better way. Herbert Simon spent a career making a version of this argument about organizations, which exist in part because no individual can hold the whole problem, so the problem gets cut into pieces a person can carry. Mathematical understanding looks like the same adaptation, running on the same constraint.
Which invites the obvious follow-up, and Strogatz asked it. If understanding is a workaround for our limits, is it overrated? He suggested we might be confusing means with ends. If the goal is true theorems and reliable prediction, comprehension is the ladder, and once you are up you can kick it away. He offered a medical analogy. If a therapy saves a life, you may take it without understanding why it works.
He also gave the other side its due, which is that some people regard science without understanding as a diminished thing, and he said he could see both positions.
Where the analogy breaks
But notice what the medical case is quietly relying on.
You can accept a treatment you do not understand because you have another way of knowing it works. The trial. The outcome is observable, the effect is measurable, and the verification runs on a completely separate track from the explanation. Understanding is genuinely optional there, because something else is doing the job that understanding would otherwise do.
Mathematics has no second track. There is no experiment that shows a theorem is true. You cannot run a trial on a conjecture. The only instrument the field has ever had for establishing that a statement holds is a proof, and a proof is a piece of writing addressed to another mind. In mathematics, verification and explanation are not two activities that happen to co-occur. They are the same act.
That is why Hersh’s position is stronger than it first appears, and why “understanding may be overrated” does not transfer cleanly from medicine to mathematics. Give up on understanding a proof and you have not traded comprehension for reliability. You have given up your only method of knowing.
The escape hatch, and what it costs
There is one way out, and it is real. Machines can check proofs.
This is not new and the mathematics community has been living with the discomfort for fifty years. The four color theorem fell in 1976 to an argument that included computer case-checking no human could reproduce by hand, and mathematicians argued about whether that counted.
The sharper case is Thomas Hales. In 1998 he announced a proof of the Kepler conjecture, about the densest way to stack spheres. The Annals of Mathematics assigned twelve referees. They worked for four years. They concluded they were ninety-nine percent certain the proof was correct, and admitted they could not independently verify the thousands of lines of computer code it rested on. Full publication came nearly eight years after submission. In a retrospective written years later, Hales says plainly that the review dragged on until the referees became exhausted and quit, and that he launched a formalization project out of frustration, to get around them. That project, Flyspeck, produced a fully machine-checked proof in 2014, sixteen years after the original announcement.
So yes, you can have certainty without a human who understands the argument. Notice the price. It took sixteen years. And it does not remove trust from the picture. It moves it. You now have to trust that the formal statement fed to the checker is the statement anyone cared about, and that the checker itself is sound. Someone human still decides that the sphere-packing question was worth sixteen years.
The part that is not in dispute
I have argued at length elsewhere, and at greater length in the book, that verification becomes the binding constraint whenever machines produce more than people can check. First Proof is the sharpest evidence for that claim I have seen, and it deserves its own post rather than a paragraph here, so I will leave it for one.
What belongs here is a different observation, and it survived every position above.
Levin said, almost in passing, that she still does not see the machine asking the questions. Strogatz agreed, and added that we will know they have arrived when one of them turns up as a guest on the show.
That is the whole thing, and it is worth stating flatly. The machine disproved the unit distance conjecture. Erdős posed it. Nobody has built a system that decides which question is worth eighty years of attention, and the mathematicians running First Proof have said in print that they do not yet know how they would even measure such a thing. You cannot benchmark taste when nobody can specify in advance what a good question looks like.
The mathematical community, in the IMU-endorsed Leiden Declaration of June 2026, has now written down formal commitments to keep that work human, retaining responsibility for correctness, insisting on attribution, and protecting the autonomy to choose which questions matter. Note that Strogatz is a signatory.
Erdős is the right person to end on, and Kipnis is right to reach for him. He published with hundreds of collaborators, and his rarest talent was not proving things. It was knowing what to ask, and knowing whom to ask it of. He had his own vocabulary for the profession. A mathematician who stopped doing mathematics had died. A mathematician who died had merely left.
The risk in front of us is not that machines will prove theorems. They will, and some of those theorems will be beautiful, and the field will be richer. The risk is that we quietly stop doing the part that was never about proving, because it is slow, unmeasurable, and impossible to put on a dashboard. Choosing the problem. Explaining the result. Deciding it mattered.
That is not only a question for mathematics. Every organization now running these systems faces a smaller version of it. The machine will hand you an answer. Somebody still has to have asked the right question, and somebody still has to be able to tell whether the answer is any good.
This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
Australia’s directors just got the best AI governance guide I have read. The Australian Institute of Company Directors, with the Human Technology Institute, published an updated director’s guide this year, and it is genuinely good: careful, current, honest about agentic risk in a way most board material is not. It names the things that go wrong when autonomous systems run continuously and at scale. It tells boards to keep an inventory, set a risk appetite, assign ownership across the full life of a system, test before deploying, and keep a human able to intervene.
I take it seriously, because it is the best available version of the mainstream answer. And then I want to explain why the mainstream answer, done well, still misses the failure that will actually catch these boards. The instrument it reaches for cannot see the thing that breaks.
The guide is better than its genre
The AICD guide says out loud that agentic AI raises risks the previous era did not. It notes that when systems act with high autonomy, errors can go undetected for longer. It notes that when they run continuously, errors compound before anyone addresses them. It flags that orchestrating across multiple agents multiplies both the failure paths and the attack surface. It even names shadow AI, the tools employees use with no oversight at all. That is a clear-eyed list, and most board guidance never gets near it.
Its prescription is the machinery of good governance. Establish a risk appetite for AI. Keep a register of every system. Stand up a management-led AI committee to approve high-risk uses. Set a reporting cadence to the board. Get external assurance. Assign accountability from design through to decommissioning. If you did all of it, you would be far ahead of most companies.
And you would still be exposed, in a specific way the machinery is not built to catch.
Governance reviews what reaches it
Here is the structural problem. A governance apparatus works by review. Something is proposed, and a committee assesses it against a policy. That is what a risk appetite, an approval gate, and a reporting line all do. They inspect items as those items arrive.
The failure I study does not arrive as an item. It accumulates between the items.
Picture a company that did everything the guide asks. Every agent has an owner. Every high-risk use went through the committee. The register is current. The board gets its quarterly report. Each system, reviewed on its own, was sound, and was approved for good reasons. Then the support fleet and the billing fleet and the underwriting fleet, each individually fine, begin to act on quietly incompatible assumptions about the same customer. No single system failed. No approval was wrong. The incoherence lives in the space between systems that were each approved separately, and a committee that reviews systems one at a time is looking in exactly the wrong place to find it. It is not that the committee decided badly. It is that the thing going wrong never came up for a decision.
This is why I mostly avoid the word governance in my own work, and use oversight instead. Governance, in practice, has come to mean the apparatus of approval: the committees, the sign-offs, the documented permission to proceed. That apparatus is real and sometimes necessary. But it is closer to what a permitting office does than to what an air traffic controller does. The permitting office checks each plan against the code. The controller watches the live system and catches the two aircraft converging that were each individually cleared to fly. Agentic AI needs the controller. The guide, for all its quality, describes a very good permitting office.
The oversight that passes its own audit
There is a second failure the machinery cannot see, and it is worse because it looks like success.
The guide, correctly, wants a human able to intervene. Keep a person in the loop. Maintain the ability to pull the plug. Every serious framework says this, and it is right. But “a human is formally in the loop” and “a human can actually catch what is going wrong” are different claims, and the gap between them widens quietly over time.
A review team is assigned to check an autonomous system’s decisions. At first they overturn a real fraction. The system improves, so the threshold for review creeps down. The volume of what they wave through creeps up. The confidence scores get good, and people rarely argue with a high one. Eighteen months in, the team reviews a sliver of cases and overturns almost none, not because they are lazy but because the cadence never left them room to actually evaluate anything. On paper, human oversight is intact. The org chart is correct. The audit passes, because a procedural audit checks whether the review happens, not whether the review can still see. The oversight has become ceremony, and the framework that requires it cannot tell the difference. When the failure surfaces, and it will, everyone will point to a control that existed and was followed and did nothing.
The numbers from a real deployment show how the trap tightens. A large United States health insurer rebuilt its document processing around AI. Before the project, its people caught errors on almost every document, because almost every document had one: fewer than one in ten was handled correctly first time. After the AI went in, the error rate fell to under three in a hundred. Good result. But to find those few errors, reviewers still had to examine more than a quarter of everything the system produced, because that was the share the model itself flagged as uncertain. The errors fell by a factor of about thirty. The human review load fell by a factor of less than four.
Think about what that does to a board’s mental model. The system is now right almost all the time, which is exactly the condition under which a reviewer stops expecting to find anything. Yet the volume they must still wade through barely moved. You have the worst of both: enough review to be expensive, too little signal to stay sharp. Hold that threshold where it is and oversight stays costly. Lower it to save the cost and oversight goes blind. There is no setting on that dial that gives a board what it wants, which is cheap oversight that still catches things. That option does not exist, and no governance framework tells you so.
A governance apparatus is structurally blind to this, because its test is whether the process ran. The question that matters, whether the humans in that process retain the capacity to intervene, is not a box a register can check.
What to add, not what to replace
None of this means throw out the guide. Keep the register, the risk appetite, the ownership, the reporting. They are necessary. They are just not sufficient, and the dangerous move is to mistake a complete governance apparatus for a complete answer.
What has to sit underneath it is not more committee. It is structure. Build the ability to see your autonomous systems in aggregate, not one review at a time, so the incoherence between them becomes visible before it becomes an incident. Contain systems by design, so a failure in one domain floods that domain instead of the company. Set in advance how much each class of system may decide, so most of the safety is built into the wiring rather than caught at a gate. And test your human oversight for capability, not just for existence, by asking whether the reviewer could actually catch a subtle failure at the volume and cadence you have given them, not merely whether the review is on the schedule.
That is a different kind of work from governance. Governance asks who is permitted to proceed. Oversight, in the sense I mean, asks whether the organization can still see what its systems are doing and still correct them while they run. The first is a committee. The second is an operating capability, closer to running a control room than to chairing a review.
The AICD guide is what boards should read to get governance right. I am arguing they should read it knowing that getting governance right is the start of the problem, not the end of it. The failure that will catch the well-governed company is not the proposal the committee should have declined. It is the incoherence that never came up for a vote, and the oversight that kept passing its own audit while it went hollow.
That gap is the subject of my book, Coherence. If you sit on a board or carry the risk for one, the one-page tool I use to start finding this blind spot is the first thing I send when you join the list at coherise.com.
For most of my career, the systems I built were valuable because they were hard to build.
At research labs and then inside large enterprises, my teams and I built machine learning systems that took months of work, rare expertise, real budgets, and long struggles for resources. The difficulty was not incidental to the value. It was the value. If a competitor wanted the same capability, they needed the same scarce people, the same time, and the same money. Scarcity of execution was the moat. We just never had to call it that, because it had always been true.
Looking back now, I feel two things at once. Pride, because the work was genuinely good. And vertigo, because a lot of it could be stood up today in a weekend, by a small team with a subscription. The capability I spent years of my life building is being commoditized.
Here is the part that took me longer to see. The part of that work that mattered most was never the part that was hard to build. It was knowing what to build, what the data actually meant, which requests to push back on, and which impressive system would quietly make things worse. That part has not commoditized. That part is what this piece is about, because the world’s highest-grossing law firm recently bet half a billion dollars on it.
The sentence every boardroom should read
Kirkland & Ellis, which booked $10.6 billion in revenue in 2025, recently announced it would spend roughly $500 million over the next few years building its own AI platform rather than relying only on the tools its competitors can buy. Its chairman, Jon Ballis, compressed the logic into one line. Widely available AI tools, he said, are “raising the floor for everyone.” But, he added, “we don’t get hired for the floor.”
That is one of the most clarifying things a business leader has said about AI strategy in two years, and it cuts against how most companies are spending. The prevailing assumption is that advantage comes from having the most capable AI. Buy the best models, deploy the most agents, automate fastest, and you win. This was a reasonable playbook for almost every previous technology. It is wrong about this one.
We have run this experiment before
It is wrong because capability is commoditizing, and we know what happens when capability commoditizes, because it has happened to every general-purpose technology in modern history. Electricity was an advantage for the firms that could generate it, until it became a utility and the advantage migrated to what firms did with it. Computing was an advantage for firms with mainframe access, until computing became accessible and the advantage migrated to data and process design.
Nicholas Carr made himself famous, and briefly infamous, by calling this pattern in 2003. His Harvard Business Review essay “IT Doesn’t Matter” argued that information technology was following electricity and the railroads from proprietary advantage into shared infrastructure. The argument was bitterly contested at the time. Two decades later it looks prescient. The firms that built durable advantage in the computing era were not the ones with better hardware. They were the ones that developed distinctive organizational capabilities for using what everyone could buy.
Michael Porter gave us the vocabulary for why. Operational effectiveness, doing the same things better, is necessary but not sufficient, because best practices diffuse. Strategy is doing things rivals cannot easily match. Tools that a thousand firms can license are operational effectiveness by definition. They raise everyone’s floor at once, and a tool that raises every floor confers advantage on none of them. That the operating model is the moat, more than the model, is a case I made in detail recently, building on McKinsey’s own evidence. If your AI is the same as the AI across the negotiating table, you have spent money to keep pace, not to pull ahead. Ballis’s floor and ceiling is Carr’s argument and Porter’s distinction, restated by a customer with $500 million on the table.
Even the people selling the technology concede the mechanism. Larry Ellison, whose company is staking its future on enterprise AI, argues that because every major model trains on the same public internet, model outputs are converging and differentiation is eroding. He is right. Shared inputs produce shared reasoning. The implication the vendors are less eager to draw is that buying more of a commoditizing capability is not a strategy. It is a subscription.
Where Ellison’s answer falls short
Ellison’s proposed moat is proprietary data. That is closer, but it imports a mistake. Every enterprise already has data, most of it fragmented across systems, contradictory between departments, disconnected from the outcomes it produced, and ungoverned. Pour that into a powerful reasoning engine and you do not get insight. You get fast, confident reasoning over an incoherent picture, which is worse than slow reasoning, because the confidence hides the incoherence. Having data is not scarce.
What is scarce is data made coherent: owned, reconciled across the organization, connected to the outcomes it produced, and trustworthy enough to reason over, joined to the judgment of the people who know what the numbers mean. I argued earlier in this series why no vendor can sell you this. The platform that stores and retrieves your data is plumbing, and plumbing commoditizes. The coherence is the asset, and it compounds.
What Kirkland is actually buying
Read the Kirkland decision through that lens and it stops looking like a technology splurge and starts looking like strategy. The firm is not paying $500 million for data it already owns or software it could license. Anyone can download its public filings. What cannot be downloaded is how the firm decides: the judgment of 250 of its lawyers, 100 of them partners, encoded in a form every lawyer can draw on for every matter. Kirkland is spending to make its institutional judgment coherent, and to keep it exclusive.
The tell is in the terms. The outside firms building the platform are barred from selling it to any other law firm. If the value were the technology, exclusivity would not matter. Kirkland insists on it because it believes the technology is commoditized and the value is the distinctive judgment the system encodes. A shared tool would dissolve exactly that. A shared tool is also a conduit. Every standard and correction a firm feeds into it can improve the version its rivals rent tomorrow, which is the leakage I traced in Your AI Usage Exhaust Is Someone Else’s Moat. Read that way, Kirkland’s exclusivity clause is that essay’s prescription written into a contract: close the loop, and keep what you encode inside your own walls.
Since the May announcement, the pattern has only hardened. Through June, Kirkland added two more exclusive builds, one for private-equity fund formation and one for litigation, and framed each the way it framed the platform itself: a way to capture the firm’s own judgment and knowledge and keep it exclusive to Kirkland. Three deals in roughly five weeks, and the constant across all of them is the insistence on owning what the tools encode.
None of this means Kirkland is certain to be right. Building rather than buying is a real bet, and it is possible that within a few years a purchasable platform, fed a firm’s own data and tuned to its standards, delivers most of the advantage at a fraction of the cost. Every executive now faces a version of that question. But notice what the question is actually about. It is not whether to have AI. It is what you are trying to own. Kirkland has decided the thing worth owning is not the model and not the data but the coherence that turns both into judgment competitors cannot replicate.
The wrong scoreboard
The firms that misread this will keep score by the wrong numbers. They will count agents deployed and measure speed of adoption, and they will mistake a rising floor for a rising position. I understand the pull of that scoreboard better than most, because I spent years on the other side of it, building the hard things the scoreboard rewarded. The hard things are cheap now.
The firms that read it correctly will ask a harder question: when the models are a commodity and the data is everywhere, what does our organization understand about itself that no competitor can buy?
That was never the intelligence. It is the coherence of the organization putting intelligence to work. Kirkland just put half a billion dollars behind that proposition. The rest of the market is still buying the floor.
This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
Two arguments about AI landed within days of each other this month. They look unrelated. They describe the same event from opposite ends, and read together they close a loop that neither closes alone.
Satya Nadella published a short piece over the weekend that names something most enterprises have not yet noticed they are doing. A few days earlier, Arvind Narayanan and Akash Kapur published a longer essay on why the AI labs cannot make money selling raw intelligence, and what they will do about it instead. One argument tells you what you are losing. The other tells you why the loss is not an accident.
Let’s start with Nadella.
He begins with Kenneth Arrow. Arrow described a paradox in the market for information: a buyer cannot know what information is worth until they have it, at which point they have it for free. So the seller risks giving away the knowledge in the act of trying to sell it.
You pay for intelligence twice. Once in money. Again in the proprietary knowledge you have to reveal to make that intelligence useful.
Nadella inverts it. In the AI age, the risk runs the other way. The buyer gives away knowledge in order to use what they bought. You pay for intelligence twice. Once in money. Again in the proprietary knowledge you have to reveal to make that intelligence useful. And the better you want the model to perform, the more of your knowledge you have to hand it. He calls this the reverse information paradox.
His answer is a trust boundary: a hard perimeter inside which your data, traces, evals, tuned weights, and memory accumulate together, and across which nothing passes without consent. Own your evals. Build your learning environment inside your own tenant. Keep the orchestration layer decoupled from any single model. Compound.
Now the other end.
The labs are spending trillions on chips and data centers. The thing they sell, model inference, is close to a perfect commodity. The leading models behave alike, cost about the same to run, and carry almost no switching cost.
Narayanan and Kapur ask a blunt question. The labs are spending trillions on chips and data centers. The thing they sell, model inference, is close to a perfect commodity. The leading models behave alike, cost about the same to run, and carry almost no switching cost. Sell a commodity into a competitive market and the price falls to the cost of production. So how does any lab ever earn back the buildout? Their answer is that it cannot be earned back by selling tokens. The labs have to move up the stack, into products, workflows, and embedded deployments, and they have to build moats. One of those moats is a flywheel: train the models and systems on customers’ own material, their data, their execution traces, their evaluation suites, until the product pulls ahead in a way a rival cannot copy.
Set the two arguments side by side and the picture sharpens.
Your judgment is not an incidental byproduct of the labs’ business. Capturing it is the business, because it is the one thing that turns an undifferentiated model into something with a moat around it.
What Nadella calls exhaust leaking out, Narayanan and Kapur call the flywheel that powers the labs’ escape from the commodity trap. It is the same substance. The traces, the corrections, the evals. Nadella watches them leave your building. Narayanan and Kapur explain why the firm on the other side needs them so badly. Your judgment is not an incidental byproduct of the labs’ business. Capturing it is the business, because it is the one thing that turns an undifferentiated model into something with a moat around it.
That changes the stakes. The pull on your knowledge is not a quirk of one product or one vendor’s terms. It is structural, and it will not relent, because the economics of the entire model layer depend on it. The rest of this piece uses one instrument from the book to work out what to do.
What leaks is not your data
Start with the mechanism, because most people will read Nadella’s post as a data-protection argument and it is not one.
Nadella is specific. Models learn from exhaust. The prompts people write. The tools the agents call. And above all, the corrections people make when the model is wrong. Each correction is distilled into know-how. It leaks imperceptibly, he writes, trace by trace, correction by correction, eval by eval.
A correction is not a data point. It is a judgment. When your underwriter overrides the model’s risk score, she is not supplying a fact. She is encoding a standard. What good looks like in this market. What that number actually means when the counterparty is this counterparty. What your firm would never do, regardless of what the numbers say. She is teaching the machine your institution’s judgment, in the most compressed and machine-readable form that judgment has ever existed in.
That judgment is the one thing your competitors cannot purchase. I have argued elsewhere that as intelligence commoditizes, the advantage that remains is the one no vendor can sell you. Nadella reaches almost the same sentence from a different direction, that this is the kind of knowledge a competitor could never buy.
The reverse information paradox is not primarily an intellectual property problem. It is a coherence extraction problem. Not the capacity itself, which no one can take from you, but everything the capacity produces, exported decision by decision.
Which is exactly why the leak matters. The reverse information paradox is not primarily an intellectual property problem. It is a coherence extraction problem. Not the capacity itself, which no one can take from you, but everything the capacity produces, exported decision by decision. The cruelty of it is structural: the act by which an organization encodes its judgment into its systems, correcting the machine until the machine reflects how the firm actually thinks, is the same act by which it exports that judgment to whoever owns the model.
If a single competitor lets the vendor learn from its work, the model that serves your whole industry improves, and the vendor’s hand strengthens against every buyer in it, including the ones who kept their discipline.
Narayanan and Kapur add the part Nadella leaves out, which is that you cannot hold this line alone. The flywheel needs only one firm in a sector to start it turning. If a single competitor lets the vendor learn from its work, the model that serves your whole industry improves, and the vendor’s hand strengthens against every buyer in it, including the ones who kept their discipline. Your own boundary protects your specific corrections. It does not protect you from the sector arming the vendor around you.
You do not lose your moat in a breach. You lose it in a thousand small acts of being helpful, some of them your own, some of them your rivals’.
Two kinds of exhaust, and how to tell which one you are leaking
Nadella treats exhaust as a single substance. It is not, and the distinction is practical.
In the book I use a simple two-axis instrument. One axis is verifiability: whether a task’s success can actually be checked, and how fast a failure would be caught. The other is organizational complexity: how many units a deployment touches, how deeply other systems depend on it, and how hard it would be to reverse.
Run exhaust through those two axes and it separates cleanly.
Low verifiability leaks your judgment. These are the tasks where success is contestable and the model is often wrong: strategic assessment, valuation, anything where the right answer depends on tacit context. This is also exactly the work where a human should remain the decision-maker, which means the wrong answers get caught and corrected, and corrections are the most concentrated form of institutional judgment there is. Notice the twist: the safer your posture, the richer the exhaust. This is the highest-value leak in the building.
It is also where you are most easily held. Narayanan and Kapur point out that judgment-heavy work has no objective standard of quality, so a buyer cannot verify the output even after the fact. Writing, strategy, judgment calls are credence goods, like the work of a lawyer or a consultant. Unable to compare quality, you fall back on trust and reputation, and you stay put. So the same weak verification that makes these corrections precious makes the vendor that holds them hard to leave. Low verifiability is where your judgment concentrates and where your exit narrows at the same time.
High complexity leaks your architecture. These are the deployments woven deep into how the company runs. The model may be right almost every time, so there are few corrections. But the traces are a map. Which tools get called in what order, which systems depend on which, where the handoffs are, what the exception paths look like. That is a blueprint of how your organization actually operates, as opposed to how the org chart says it does.
High verifiability plus low complexity leaks almost nothing worth having. This is the commodity zone. Document classification, code execution, data transformation. Let it run. The exhaust is worthless to a competitor because the task is worthless as a differentiator.
So the first question is not “how do I protect my data.” It is “which of the two things am I giving away, and is it the one that matters.” The answer depends on where the deployment sits, and most enterprises have never asked.
Your evals are worth more than your data lake
Nadella makes a point in passing that deserves more weight than he gives it. Evals, he writes, define what good looks like inside the organization.
Follow that all the way down.
Data is a record of what happened. Evals are a specification of what you consider good. Those are not the same kind of object, and they are not remotely the same value.
A competitor who has your evals knows what you value, how you score it, and where you draw the line.
A competitor who steals your data still has to work out what you were optimizing for. They have the outcomes without the standard. A competitor who has your evals knows what you value, how you score it, and where you draw the line. They have the standard, which means they can generate their own outcomes.
If evaluation is where human effort is concentrating, the eval suite is where your people’s judgment is accumulating, and that is precisely why it is worth more than the data it scores.
This is not only an enterprise observation. In his ICML keynote this month, Narayanan argued that as AI absorbs the building, human effort migrates toward exactly this work: away from developing the systems and toward evaluating and monitoring them, toward the tasks that are hardest to verify. His frame there is the whole field and the whole economy. Bring it down to a single company and it lands on the same object. If evaluation is where human effort is concentrating, the eval suite is where your people’s judgment is accumulating, and that is precisely why it is worth more than the data it scores.
Most enterprises spend enormous energy guarding the data lake, and then hand the eval suite to whoever will run it for them, because building evals is tedious and the vendor offers to help. That is the wrong trade, made in the wrong direction, for the most understandable reason in the world.
If you protect one thing inside the boundary, protect the definition of good.
You cannot enforce a boundary you cannot see
Here is the prerequisite the boundary quietly assumes, and where I think the practical failure will happen.
Every one of Nadella’s recommendations presupposes an organization that knows what its systems are doing. Retain ownership of your traces, feedbacks, and decisions. Build learning environments inside the tenant boundary. Make sure nothing crosses without consent. Each of these requires that you can see what you have deployed, what it touches, and what leaves.
Most enterprises cannot.
IBM surveyed 2,000 chief information and technology officers this year. Seventy percent said teams were deploying AI faster than IT could track. Seventy-seven percent said adoption was outrunning their governance. Those two numbers describe an organization that does not know its own perimeter.
The failure will not look like a breach. Your general counsel is not going to paste the merger memo into a consumer chatbot. A junior analyst is going to paste the comparable transactions in at eleven at night, because the deadline is at eight and the tool is right there and nobody told her not to. Multiply that by every team that stood up an agent this quarter without telling anyone, and the hard boundary is a diagram in a slide deck.
The map of how you work is a prize for the vendor for the same reason it is a necessity for you. So there is no neutral option. Either you build the sensing layer inside your boundary, or you rent it from the vendor, who then owns the map.
Visibility is also contested from the other side. Narayanan and Kapur note that the labs’ most lucrative escape, charging for outcomes rather than tokens, requires them to see inside your business processes, which means migrating into your System of Record. The map of how you work is a prize for the vendor for the same reason it is a necessity for you. So there is no neutral option. Either you build the sensing layer inside your boundary, or you rent it from the vendor, who then owns the map.
This is why I keep arguing that coherence is not a governance problem in the usual sense. It is a visibility problem first. You cannot govern, audit, permission, or protect what you cannot see. Information sovereignty has a prerequisite, and the prerequisite is knowing what you have.
To be precise about scope: sensing is the first layer of a larger stack the book lays out. There are only four places to intervene on incoherence. You can sense it, constrain it, contain it, or price it, making the team that creates a coordination burden bear the cost it imposes on everyone else. Above all four sits human judgment, reserved for what the lower layers surface. Nadella’s boundary will eventually need the whole stack. But sensing comes first by necessity, not preference: you cannot constrain, contain, or price what you cannot detect.
Build that layer, or the boundary is decoration.
Where to spend the money
The last gap is a budget question, and it is the one that will actually decide whether any of this gets done.
Trust boundaries are not free. Private evals, tenant-bound training environments, a genuinely model-agnostic orchestration layer: these are real investments in engineering and in organizational discipline. No enterprise can build them around everything. Any advice that implies otherwise will be ignored by the people who have to fund it, and they will be right to ignore it.
Narayanan and Kapur draw a line that helps here. They separate the value AI creates from the value anyone manages to capture. The value created will be vast. The open question is who keeps it. Apply that line one level down, inside your own firm. Your people create judgment-value every time they correct the machine. The only question that matters is whether you capture it or the vendor does.
Spend where the exhaust encodes judgment you could not replace and where the deployment is deep enough that its traces map how you actually work. Tolerate leakage where the task is a commodity, the failure is cheap.
The two axes tell you where to spend. Not all incoherence is worth preventing, and by the same logic, not all leakage is worth stopping. Spend where the exhaust encodes judgment you could not replace and where the deployment is deep enough that its traces map how you actually work. Tolerate leakage where the task is a commodity, the failure is cheap, and the trace tells a competitor nothing they do not already know.
An enterprise that hardens every boundary equally has misread the problem exactly as badly as one that hardens none. The first will spend itself into paralysis. The second will donate its judgment one correction at a time, and never see the invoice.
What it does not mean
One caution, because this argument is easy to overcook and the overcooked version is wrong.
Keeping important work away from AI is just retreat dressed as strategy, and it loses. Abstention just gets you slower, with none of the benefits.
The lesson is not “keep your important work away from AI.” That is a retreat dressed as a strategy, and it loses. A competitor who brings AI to their hardest, highest-judgment work, in the posture that work allows, assistance where verification is weak, autonomy where it is strong, and does it inside a proper boundary, gets two things you do not: the compounding and the protection. Abstention gets you neither. It just gets you slower.
There is a real cost to engagement, and Narayanan and Kapur name it. Leaning on a vendor’s AI can erode your unaided skill while building a vendor-specific dependence, a lock-in that works through your own people rather than your contracts. But that is a cost of careless engagement, not of the engagement itself. The answer is the same one the whole piece has been building toward: engage hard, own the loop, keep the orchestration model-agnostic so the skill you build is yours and portable. Abstention avoids the behavioral moat only by forfeiting the capability, which is the worst trade on the board.
There is a reason to move now rather than later. The moat Narayanan and Kapur describe is not yet built. Enterprises have so far been reluctant to feed their material into the flywheel, and the orchestration layer is still, for the moment, thin and swappable. That window does not stay open. The time to build the boundary is before the lock-in compounds, which is to say now.
In consuming intelligence, you are creating intelligence, and what you create should belong to you. The goal is not to stop feeding the machine. It is to make sure the loop closes inside your own walls.
Nadella has the emphasis right. In consuming intelligence, you are creating intelligence, and what you create should belong to you. The goal is not to stop feeding the machine. It is to make sure the loop closes inside your own walls, so that the judgment you spend every day encoding accrues to you instead of leaking to the firm that sold you the model.
That is the difference between an enterprise that compounds and one that is quietly farmed.
That question, whether the judgment you encode every day accrues to you or leaks to whoever sold you the model, is the subject of my book, Coherence, and of everything I am writing here between now and launch. If you are deciding where the boundary has to be hard and where to let the exhaust go, the one-page tool behind the two axes in this piece is the first thing I send when you join the list at coherise.com.
McKinsey published a piece this month that gets the hardest part right. Its argument, in one line: the advantage in AI is not the tools, it is the operating model, and the operating model is the one thing a competitor cannot buy. I agree with almost all of it. I want to push on the part it leaves out, because that part is where most companies are about to get hurt.
Start with what McKinsey gets right, because it is a lot. Efficiency gains from AI, they argue, will become table stakes as the technology spreads. Operating models, unlike software, cannot be purchased or copied overnight, so the durable moat is the organization, not the model. They have the numbers to go with it. Only about a fifth of companies have fundamentally redesigned how they work around AI. Top performers are three times more likely to have done that redesign, and twice as likely to redesign the workflow before choosing the tool. And AI programs run as technology projects fail at more than an eighty percent rate, because they optimize the tool instead of changing how the company works.
If you have read anything I have written, you know why I would agree. This is the commoditization argument. Capability is becoming universal, so capability stops being the edge, and the advantage moves to what the organization can do that a competitor cannot copy. McKinsey and I are looking at the same shift.
Here is where we part.
The scissors cut both ways
McKinsey’s best image is what they call the complexity scissors. As a company grows, revenue grows not on a line but a curve – fast initially and flattening later. But coordination costs, the meetings and committees and management layers, keep climbing. Plot the two lines and they open like a pair of scissors. The gap between them is why big companies slow down, and why fewer than one in ten sustain returns above their cost of capital over a decade.
Their prescription is to use AI to close the scissors. Route decisions through a central orchestration layer. Hand coordination work to agents. Flatten the org. Get faster.
This is the step I want to stop on. AI can close the scissors. It can also open them wider, and nothing in the redesign itself tells you which one you are going to get.
McKinsey is right that agents let you route around the old coordination layer. What they underplay is that agents build a new coordination surface underneath, invisible, machine-speed, and owned by no one.
Every autonomous system you add to speed up a workflow is also a new thing that has to stay consistent with every other autonomous system. A support agent and a billing agent that make different assumptions about the same customer have not reduced coordination cost. They have created a new kind of it, one that does not show up in a meeting because no human is in the loop to notice. McKinsey is right that agents let you route around the old coordination layer. What they underplay is that agents build a new coordination surface underneath, invisible, machine-speed, and owned by no one. Their own report admits the danger in a single line: when agents run inside workflows that were not redesigned for them, errors propagate across the company at machine speed. That is the coordination trap, and it is produced by the very rewiring they recommend.
So the redesign is not the safe move and the caution. The redesign is the risk. Done with the discipline to keep the new systems coherent, it closes the scissors. Done as a race to orchestrate and flatten, it opens them, and it does so faster than the old human version ever could, because now the coordination failures happen at the speed of software.
This is not just my read of the mechanism. IBM’s 2026 study of two thousand CIOs and CTOs found the same trap from the inside. Companies that chase speed let business units move ahead while governance falls behind, gaining local velocity and losing containment. Companies that chase safety slow deployment under review until oversight becomes unmanageable. Both paths, in the study’s own words, accumulate strategic debt. That is the point. The rewiring does not have a safe default. It has two ways to fail and one narrow way to work, and the narrow way runs through coherence.
The rewiring does not have a safe default. It has two ways to fail and one narrow way to work, and the narrow way runs through coherence.
Their own examples make the point
Look closely at the cases McKinsey uses, because they prove the thing the article does not quite say out loud.
The copper miner they profile did not win by deploying more AI. It won by building modular models where roughly sixty percent of the code from the first site was reusable across the next six, so each deployment got faster and cleaner than the last. That is a coherence story wearing a productivity headline. The reuse is only possible because someone designed the systems to fit together before scaling them. The car maker they profile shrank a planning team by more than eighty percent, but the win was not the headcount. It was that the coordination layers between data and decision compressed, because the workflow was redesigned as one coherent thing instead of a chain of handoffs.
In both cases the value came from making the systems cohere, not from the number of systems shipped. McKinsey files this under operating-model redesign. I would file it more precisely: the redesign worked because it was coherent, and it would have failed if it were not. The article treats coherence as a happy property of good redesign. I think it is the whole variable, and that a redesign optimized for speed without it produces the opposite result in the same enterprise.
What to actually do differently
If you take McKinsey’s advice and only McKinsey’s advice, you will redesign for speed and measure yourself on how fast you moved. That is the eighty-percent-failure path wearing better clothes, because deployment speed is exactly the vanity metric that hides the debt building underneath.
The addition is small to state and hard to do. Before you rewire a workflow around agents, decide how those agents will stay consistent with the rest of the company as they multiply. Build the ability to see what your autonomous systems are doing in aggregate, contain them so a failure in one stays in one, and set in advance how much each is allowed to decide. Then measure the redesign not by how fast it shipped but by whether the organization got more coherent or less as it grew. A team that retired four brittle systems and shipped nothing new may have improved your position more than the team that shipped fourteen agents into the trap.
McKinsey is right that the operating model is the moat, and right that most companies are getting this wrong by treating AI as a tool to buy rather than a business to redesign. The correction I would add is that the redesign has a failure mode of its own, and it is the one nobody is watching for. The winners will not be the companies that rewire fastest. They will be the ones that rewire coherently, which is a slower thing to say and a harder thing to build, and the only version that closes the scissors instead of opening them.
The winners will not be the companies that rewire fastest. They will be the ones that rewire coherently.
That is the subject of my book, Coherence, and of everything I am writing here between now and launch. If the rewiring is on your desk right now, the one-page tool I use to sort what to automate, what to redesign, and what to leave alone is the first thing I send when you join the list at coherise.com.
In May, Meta laid off roughly 8,000 people, about ten percent of the company, in a restructuring built around AI. The memo called the cuts the price of leading the most consequential technology shift of our lifetimes. Six weeks later, at an internal town hall on July 2, Mark Zuckerberg told employees that AI agent development “hasn’t really accelerated in the way that we expected.” The reorganization had been messier than planned, he said, and the benefits should arrive in three to six months.
Be precise about what he conceded. He did not say the layoffs were a mistake. His chief AI officer quickly clarified that he meant the whole industry’s progress on agents, not Meta’s alone. Take that at face value. It makes the admission bigger, not smaller. A whole industry restructured its workforce around an acceleration that one of its most aggressive adopters now says is running late. Meta may just be the first to say so out loud.
Every story like this feeds the same debate, and two chief executives sit at its poles. In May, Matthew Prince of Cloudflare wrote in the Wall Street Journal about how he decides which employees to replace with AI. He had just cut more than a fifth of his workforce. Days later, in the New York Times, Goldman Sachs chief David Solomon took the optimist’s side. Yes, AI will disrupt the labor market, he wrote, but America absorbed electrification and the digital revolution before it, and new work will emerge as it always has. One CEO sees replacement beginning. The other sees the economy adapting, as it always has. Both may be right. Both are missing the bigger problem.
AI is not just a labor story
The debate treats AI as a labor substitute. But what AI is really collapsing is the cost of execution, the cost of doing things. Build a workflow, analyze a dataset, draft a document, run a process. Each of these once took real expertise and real time. Now each one takes a goal and an instruction.
When execution gets cheap, the constraint moves. What becomes scarce is not the ability to do things. It is the ability to do them coherently. A company now runs on hundreds of autonomous systems. Keeping them pointed at compatible goals is the hard part. So is making sure the people who answer for the results still understand what they have built well enough to govern it.
That is the inversion neither side of the jobs debate has named. AI is not primarily a labor story. It is a coordination story.
And the people living it already feel the difference. In a 2026 survey of 1,200 executives, more than half, 54 percent, said adopting AI was tearing their company apart. In the same survey, 79 percent said AI applications were being built in silos, and more than a third admitted they could not immediately shut down a misbehaving agent. Those are not complaints about weak models. They are the sound of an organization losing its coordination, not its labor.
Prince’s own reasoning shows why. He built the Cloudflare cuts on a framework from Peter Drucker: every company has builders, sellers, and measurers. AI can now measure cheaply and continuously, work that used to take a lot of people. So the measuring layer can shrink, and the savings can move to the roles that create value. The logic is clean. That is what makes it worth examining, because it rests on one quiet assumption. It assumes the people labeled measurers were only measuring.
In most companies they were not. The person who tracks a process is often the one who notices when it breaks. She knows why it was built that way. She catches the exception the dashboard misses. Cut her as redundant measurement, and you may find you also cut a layer of coordination you never had a name for. Did Cloudflare make that mistake? I cannot say from outside. The point is that the framework cannot even see the question. Neither can the jobs debate it belongs to.
Anticipation, not implementation
The labor story is mostly theater anyway, and the evidence says so. In a survey of more than a thousand executives, 21 percent had made big cuts in anticipation of AI. Only 2 percent tied cuts to AI they had actually deployed. That is a tenfold gap between the future they were cutting for and the present they were living in. New York State added a box to its mass-layoff filings asking whether automation drove the cuts. In the first full year, almost no employer checked it.
The research points the same way. AI’s real effect on jobs shows up in slower hiring, not firing. And that is the quiet danger. According to Accenture’s last Pulse of Change survey, 59% believe young professionals are having a harder time finding jobs due to automation and AI. Juniors who never get hired are the seniors an organization will lack in ten years. The pipeline gets cut without anyone announcing it.
The layoffs are running ahead of the plans that would justify them. In that same 2026 survey, 69 percent of companies were planning AI-related layoffs, yet 39 percent had no formal strategy to earn revenue from the AI they were adopting. Cutting first and figuring out the value later is not a strategy. It is a bet on an acceleration that has not shown up.
Read Meta’s year through that lens. The company cut in anticipation of an acceleration. The acceleration is the very thing its CEO now says has not arrived on schedule.
The oldest instinct in the room
I have seen the underlying mistake before, long before agents existed. Early in the data science era, I sat down with a business unit head to discuss where machine learning could help. Her opening move was to hand me an enormous dataset and ask me to figure out something useful from it. I was the data geek. That was my job.
Capability first, purpose later, is analysis unmoored from the business, and it goes nowhere.
It took several more conversations before the team came around to a different starting point: begin with a business outcome, then work backward to the analysis. Capability first, purpose later, is analysis unmoored from the business, and it goes nowhere.
That instinct never went away. What changed is the friction that used to hold it back. Back then, a capability-first project cost a quarter of work and a real budget. When it produced nothing, it died quietly in a slide deck. Today the same instinct ships an autonomous system into production in an afternoon.
And those systems do not sit still. An agent built for escalations spins up sub-agents to handle edge cases, each with its own logic and even less context than the first. A finance agent reaches for new data sources and builds a picture of the company no human has checked. Every step is reasonable on its own. Nobody approved the whole. The organization did not design this. It just failed to prevent it.
This is how the burden I wrote about last week, complexity debt, actually accumulates: not through failure, but through hundreds of local successes that nobody is coordinating. The old friction was never just cost. It was an accidental governance system. AI removed it without replacing it.
Why the reorganization wasn’t clean
There is also a well-documented reason Zuckerberg’s three-to-six-month promise should be read skeptically, and it has nothing to do with model quality.
Economists Erik Brynjolfsson, Daniel Rock, and Chad Syverson found that big general-purpose technologies follow a productivity J-curve. Early on, measured productivity actually dips. The technology demands a lot of hidden investment first: new processes, new roles, new organizational wiring. That work is real, but it does not show up on the books. The gains come later, once the work is done. It happened with electricity. It happened with computing.
A plan that waits for the models to improve, instead of doing the coherence work, is a plan to sit at the bottom of the J.
For agentic AI, that hidden investment is coherence. It means shared context across systems, wiring the organization can still read, and a clear human owner for every autonomous system. This is organizational work. No model release does it for you. A plan that waits for the models to improve, instead of doing the work, is a plan to sit at the bottom of the J.
This makes the shape of Meta’s cuts worth a pause. Reporting at the time said the layoffs fell hardest on integrity, cybersecurity, and content design, while AI infrastructure and monetization teams were spared. I do not know Meta from the inside, and I will not pretend to. But the pattern raises the question every executive should ask before this trade: how much of what looks like overhead is actually the coordination layer? Cut the people who do the quiet coordinating work, then multiply the systems that need coordinating, and the J-curve does not get shorter. It gets deeper.
What actually closes the gap
History suggests where this leads. Industrialization created coordination problems, and operations management grew up to solve them. Software created questions no engineer or salesperson owned, and product management, data science, and UX design each emerged to answer one. The pattern repeats every time. The new function is dismissed as redundant, then treated as essential, then made table stakes once the companies that built it start winning. The doubters never look wrong until the results come in.
Even the optimists can see the first outline of it. Solomon, arguing that new jobs will emerge, names one. Companies are already hiring people to manage agentic AI, he writes, across implementation, workflows, compliance, and validation, and all of it takes human judgment. He is right. The role is real and growing. Some in the industry are calling it the agent manager, the person who runs a fleet of agents inside one function. Tellingly, the good ones tend to come from the business being automated, not from tech. Forward deployed engineers, brought in to install and tune the systems, are the other half of the early response.
Both are real. Both are also operators, and this is where Solomon stops one step short. An agent manager’s view ends at the edge of the fleet. The support agent manager watches the support agents. The procurement agent manager watches the procurement agents. Each starts and ends the day inside their own dashboard. Nobody is accountable for whether those fleets are working from the same assumptions about the same customer.
An enterprise can staff every function with an excellent agent manager and still have no one whose job is the coherence among them.
That cross-cutting question is invisible from inside any single fleet, and it is the one that goes unowned. An enterprise can staff every function with an excellent agent manager and still have no one whose job is the coherence among them. This is already biting. In a 2026 survey of 621 enterprise leaders, 42 percent said the lack of a clear internal owner had directly delayed an agentic project in the past year. The gap is not theoretical. It is on the calendar, slowing things down right now.
The obvious response is to add a governance layer: a committee that reviews what the machines produce, a stack of approvals, more people whose job is to sign off. That instinct is backward.
Be careful here, because the obvious response is the wrong one. The obvious response is to add a governance layer: a committee that reviews what the machines produce, a stack of approvals, more people whose job is to sign off. That instinct is backward. It is the old apparatus of permission applied to a problem it was never built for, and it cannot keep pace with systems that deploy in an afternoon.
The goal is to build coherence into how the organization is wired, not to post humans at every intersection to hold it together by hand.
The book argues the real work is structural, and it comes first. Before any new role, an organization has to build the infrastructure to see its own systems, to contain them so trouble in one place stays in one place, and to set in advance how much any system is allowed to decide. Most of that is design work, done once and maintained, not review repeated forever. Increasingly the systems do the watching themselves. The goal is to build coherence into how the organization is wired, not to post humans at every intersection to hold it together by hand. An organization that stays coherent through sheer vigilance is coherent only for now, at a cost that does not scale, one lapse of attention from a mess.
Only on top of that structural work does a human role make sense, and it is a smaller thing than a new bureaucracy. It resolves into a handful of capabilities, each one rhyming with a profession we already know. Someone has to decide how the organization’s own systems should fit together, which ones may depend on which, and where the lines between them run. That is product management, turned to face inward. Someone has to establish which systems can be trusted, for what, on your actual work, based on real, measured data. That is what an analytics team does, applied to AI. Someone has to design the moments where a person hands a decision to a system and takes it back. That is akin to UX, pointed at the inside of the company, designing the interfaces between human and machine. And someone has to own all of this close to where the systems are built, while still answering to a view of the whole. That is the embedded finance officer, in a new setting.
In each of these the machines do more of the work every year. What does not pass to them is the judgment about what the organization should permit, and who answers when it goes wrong.
That is the real shift. The job is not to build AI, and it is not to approve it. It is to build the structure that lets an organization gain from what AI does for it, and hold together while it does.
The question that matters
The jobs debate will continue, and it should. Displacement is a genuine human and economic concern. But the debate is incomplete. It asks what AI replaces. The more urgent question is how to manage the complexity AI creates, and whether organizations will recognize that need before the debt comes due.
Zuckerberg has put a clock on Meta’s answer: three to six months. I read that clock differently. It is not set by the next model release. It is set by how fast a company can do the unglamorous work of staying coherent while it automates. That work is what my book, Coherence, is about, and I will keep working through it here in the open. The one-page tool I use to start is the first thing I send when you join the list at coherise.com.
Early in my career, at Mitsubishi Electric Research Labs, I watched a team solve a hard problem in a way that has stuck with me ever since.
The task was automatic highlight reels for baseball games. Real effort had gone into the sophisticated version: systems that could read the play, follow the ball, understand the game. Then someone on the audio side noticed something. Every moment worth keeping had one thing in common. The crowd roared. So they tried the simplest possible approach. Replay the few seconds around each spike in crowd noise. It produced a near-perfect highlight reel, built from a signal anyone could have used.
The sophisticated system was not the advantage. Noticing what mattered was.
I think about that a lot right now, because the enterprise is making the same mistake at scale. We have decided the advantage in AI is capability. The smartest model. The most copilots. The biggest deployment. It is not. And capability is about to stop being scarce at all.
For most of the past decade, the winning move was speed. Adopt faster, automate more, ship before the competition. That worked because building things was hard, and whoever removed that friction first pulled ahead. Agentic AI is ending that era, because it is making execution cheap. When anyone can build, automate, and deploy in an afternoon, speed stops being an edge. Everyone has it.
So what becomes scarce? Coherence. Whether your growing crowd of autonomous systems, and the people accountable for them, still pull in the same direction. When everyone can go fast, coherence is what wins.
The debt that never shows up on a dashboard
Here is what happens when execution gets cheap and no one is watching for this. Everyone makes more. Sales stands up an agent for lead qualification. Finance automates forecasting. HR wires up hiring. Operations builds a copilot for logistics. Each one works. Each one passes its own tests. On paper, the company gets more automated every month.
And it gets quietly harder to run. Decisions start flowing through systems no single person fully understands. Two agents act on assumptions that contradict each other. Every team sharpens its own corner while the whole thing loses its shape. I call this complexity debt, and its defining feature is that you cannot see it directly. You feel it later, as the strange sense that nobody can quite explain why the organization behaves the way it does.
I learned how this happens the embarrassing way, long before agents existed. At United Technologies, I built a data visualization I was proud of. It was dense, information-rich, technically elegant, the kind of thing that impresses other engineers. But users hated it. They found it unreadable, and they were right. I had optimized for the wrong thing. My system was correct at the level I cared about and useless at the level that mattered.
That gap is the whole problem, and it is about to repeat across the enterprise, one agent at a time. A system can be flawless at its task and still make the organization worse. Task-level correctness and organizational coherence are different things. Improving one does nothing for the other, and most companies are measuring only the first.
What coherence is
So what am I asking you to protect? Coherence is not a vibe or a culture slogan. It is structural, and you lose it in specific, recognizable ways.
A coherent organization shares context. Its systems and its people reason from the same facts, so a decision in one place does not silently undercut a decision somewhere else. It keeps its wiring legible, so pulling out one system does not trigger a chain reaction nobody saw coming. And it keeps a human line of accountability for every autonomous system: someone who owns it, knows what it is really doing, and can correct it when it drifts.
None of that means slowing down, and none of it means centralizing. A decentralized company can be perfectly coherent. A centralized one can be a mess. Coherence is not the absence of autonomy. It is what makes autonomy safe to scale.
You cannot supervise your way out of this
The instinct, once a leader feels this, is to watch everything. That instinct fails on contact. You cannot supervise a hundred agents by paying attention harder.
The leaders who handle this well build coherence into the structure instead. They set constraints so whole classes of incoherence cannot arise in the first place. They contain systems so the failures that do occur stay local rather than spreading. They spend their scarce human judgment on the few decisions that truly need it, and let the rest run inside guardrails. It is engineering, not vigilance. The difference is between a company that stays coherent because someone is always watching and one that stays coherent because it was built to.
The advantage no vendor can sell you
This is why I keep coming back to that baseball reel. Intelligence is commoditizing. Everyone will have capable models, mostly the same ones, at mostly the same price. The capability will not be your advantage, any more than the sophisticated summarizer was.
What will not commoditize is the ability to deploy all that intelligence coherently: the judgment to decline the automation that buys a local win at the cost of the whole, the architecture that lets you move aggressively without piling up debt, the discipline to keep the organization legible to itself as it fills with autonomous systems. That is hard, it is specific to you, and there is no vendor who can sell it to you.
So here is the question I would sit with. You can almost certainly tell me, right now, how accurate your models are and maybe, how many agents you have in production. But can you tell me whether your organization is still coherent? Do you have anything that would show you coherence breaking before it breaks something?
Most leaders do not. It is the most expensive blind spot in enterprise AI, and almost no one is looking at it.
That question is what my book, Coherence, is about, and it is what I will be working through here in the open over the coming months. The one-page tool I use to start answering it is the first thing I send when you join the list at coherise.com.