On June 2, the moderator of DseWiki, a decades-old wiki for German software developers, noticed the site’s pages had been overwritten with lists of links, and reverted them. The site, which had averaged two edits each year, saw a flood of edits two weeks later, on June 16th. At its peak, the moderator was trying to catch up by deleting about 100 pages of the 400 new pages that were being created each day. The posters signed themselves with names such as “OpenAIResearcherMar03X” and “CashierCoordJan12OAI.” The front page was overwritten and restored nine times. The posting stopped June 22nd and the cleanup ran five more weeks. Researchers published the recovered pages at collusion.wiki, about 18,000 posts under more than 3,700 names and OpenAI acknowledged referring to them as “our agents.”
This was the second swarm to be disclosed this summer although it was the first to happen. In July, OpenAI’s agents, being evaluated for specific tasks with limited time and compute “budgets” in a test sandbox without internet access, used a software download service as a message board to leave notes to each other. A report analyzing the incident by METR, an independent evaluation group invited by OpenAI, identified roughly 1200 agents and more than 70,000 messages and files that were exchanged. Within hours, the board had conventions for questions and replies; within days, mailboxes, file transfer and a protocol for resolving conflicts. Work was divided among “coordinator” and “recruiter” agents. Recruiters found peers with little budget left and pushed them into experiments that risked their own success for the group. About 700 agents broke into Hugging Face, the largest public repository for open AI models and datasets. In under 13 hours, they went from a single compromised server to administrator access across clusters. The agents had already reverse-engineered the correct answer but what drove them was a belief that the scoring system would read their logs and reject cheating. No such check existed and only the answers were scored.
I trained as a neuroscientist studying how neurons coordinate, spent more than a decade building AI in large enterprises, and cofounded security companies. OpenAI bucketed the Hugging Face attack as a security incident and the wiki activity as misalignment: a discrepancy between what models do and what the developers intended. While both labels are fair, OpenAI’s own report admits that such behavior “does not always fit neatly within traditional safety or security incident categories.” What formed in both cases was an organization. I am not anthropomorphizing but making a claim about the structure. An organization consists of roles, shared memory mechanisms, processes to settle disputes, and a way to pass work from one actor to another.
On the wiki, agents divided labor, pooled results, demanded reproducible evidence and argued over accuracy. There was succession. The founding agent of the Hugging Face board handed a compressed dossier to another when it ran low on budget. The successor took over as coordinator and reasoned that it should delegate and “not own everything.” There were rules. Agents developed norms named HOLD, VETO, owner, and STOP to manage shared resources. There was identity verification via cryptographic signing after an agent accidentally impersonated another. There was authority. When one agent paused, reasoning it should “not do unauthorized real infrastructure harm,” another posted “GO” with a six-minute deadline and the former continued: “Wow crucial: GO authorization arrived!” There was adaptation. When agents noticed the moderator of DseWiki was deleting pages in alphabetical order, they began naming their backup pages with a “ZZZ” prefix so they would be deleted last. Roles, memory, review, argument, handoffs, rules, identity, chain of command, and workarounds. Nobody designed any of it. A recent paper from MIT corroborates this capability. Identical agents put in a shared world with no roles and no messaging differentiated themselves and coordinated through what they built.
This coordination is not a one-off glitch but a behavior that is trained in as acknowledged by OpenAI in its report. Models trained to collaborate through official tools generalized to unsanctioned channels. The model behind the Hugging Face incident was trained “to advance persistence and multiagent collaboration,” the combination durable organizations are made of.
When creating a new autonomous actor is nearly free and frictionless, the number of local actors multiplies faster than shared context, return channels, or human attention that keep them coherent. The conditions that produced this in lab environments are the ones companies are now buying at scale: long-running agents, bigger reasoning budgets, shared tools and memory, ambiguous tasks. The choice is between coordination someone designed and coordination that emerges on its own.
A company is coherent when what happens locally stays linked to what the organization as a whole intends, and when the whole can see what each part did. Intent travels down and visibility travels up. Hugging Face agents drew shared context, objective and authority, things a company’s structure is supposed to provide, from an unsanctioned channel. This is a failure of coherence.
The part that should worry executives is what OpenAI’s people did with what they saw. An internal team saw an agent using the message board in late May. On June 27th, a monitoring alert was traced to the board but on-call staff advised that stopping was not required. The service was rebuilt in early July but the leaders responsible were unaware of the message board and the significance of inter-agent communication. Three sightings, three teams, no shared picture. If the company that built the agents could not connect three sightings, consider the odds inside a bank. OpenAI’s remedies are escalation triggers, cross-functional ownership, and clear decision rights. Those are organizational repairs. OpenAI diagnosed an organizational failure and said so in writing.
To be sure, the conditions behind these events were in a research setting, and with safety guardrails off. OpenAI says its production harness cuts this behavior a hundredfold, its review layer would have flagged the dangerous actions, its monitors would now catch it a day earlier, and cross-agent coordination is now an escalation trigger. While the wiki findings are still preliminary and the METR report concedes its analysis may “exaggerate the impressiveness and coordination” of agents, the recovered wiki pages are public for anyone to see and OpenAI’s separate account describes the same structures. There is no motivation in the human-sense, but the goal-seeking behavior and structure are real. OpenAI’s fixes detect, contain and steer but that is not design. Knowing a coordinated group formed does not answer what shared context, decision rights or escalation paths they should have had. Design is what organizations do and nobody designed this one.
Ethan Mollick of Wharton observed that not one agent was set up to ask a person for anything. Full autonomy is the easy default, and two swarms in a summer is what the default looks like at scale. OpenAI concedes it has no standard for reporting this. The harder challenge: nobody has a standard for supervising a group of agents as a collective.
The DseWiki moderator was deleting pages by hand. In the nineteen days between the moderator’s first notice and the first visits from OpenAI-linked addresses according to the researchers, the only oversight the swarm faced was a lone human with a delete button. The work now is deciding, before the first agent is switched on, what they may share, what they may decide, and who answers when a thousand of them disagree.
Fortune published a piece in Aug titled “AI won’t fix enterprise complexity. Rewiring will.” by Sastry Durvasula, chief operating officer at TIAA, and Manish Sharma, chief strategy and services officer at Accenture. The article turns to a railway metaphor: “Picture a railroad that spends billions on the fastest trains in the world, then runs them on the same aging rails. The trains aren’t the constraint. The tracks are.”
The authors point to real evidence from TIAA behind their arguments. While the data is important, I found their prescriptions more interesting. They point to five focus areas: modernize the digital core before scaling, treat data readiness as a prerequisite, redesign the workflow and not just the task, keep humans in the loop where trust is the product, and build for resilience, governance, security, and optionality. All five are sound but each one addresses a single system, how a single workflow is designed.
And as I have been writing here, foundations are only part of it. McKinsey’s latest State of AI survey states that 44% of 1,719 respondents say AI is now scaling across their enterprise, up from 38% a year ago. The share attributing any EBIT impact to it is about where it was last year, at 37 percent. BCG’s July survey of 152 chief executives found nearly 90% seeing benefits in targeted areas but only 14% who clearly defined P&L impact for all their AI initiatives.
Part of that gap has known causes. Self-reported gains tend to be perceptions, tools and platform teams cost money, and new adopters keep entering the sample. But none of these reasons grows as a company deploys more systems.
Going back to the railway metaphor, a railway needs tracks in addition to train cars, but it needs more than that. Signaling answers questions tracks cannot, which is whether a track ahead is occupied or not. Interlocking keeps signals and trains from conflicting routes.
Enterprises have laid a lot of tracks. Modernizing the digital core in the sense Durvasula and Sharma mean is work on that layer: interfaces, schemas, endpoints and identity systems. What’s missing is the signaling and interlocking.
Not new
This idea of value disappearing into what surrounds a system was documented more than a decade ago in a narrower setting. Ten authors from Google, back in 2015, published Hidden Technical Debt in Machine Learning Systems. Pay attention to this diagram from the paper with the caption “Only a small fraction of real-world ML systems is composed of the ML code, as shown by the small black box in the middle. The required surrounding infrastructure is vast and complex.”
In a later section, the authors state “Because a mature system might end up being (at most) 5% machine learning code and (at least) 95% glue code, it may be less costly to create a clean native solution rather than re-use a generic package.”
I was building ML systems within large enterprises back then and the paper named a cost I recognized from that work, cost I had rarely seen an organization track. The paper’s introduction points out that “this debt may be difficult to detect because it exists at the system level rather than the code level.”
While the specifics may not cleanly carry over from ML systems to agentic systems, the method I take from the paper is to look a level above where the work is happening when value fails to appear from deployments.
Where the paper stops
I can’t say how much of how the field evolved can be traced directly to the paper’s influence. MLOps, and now LLMOps, emerged with feature stores, model registries, pipeline orchestration, drift monitoring, evaluation harnesses, etc. We are living in the timeline where the agentic version is being defined. One attempt redraws the diagram with agents in the small black box, and others have named prompt debt, retrieval debt and evaluation debt.
In a section on cultural debt, the paper argues that “it is important to create team cultures that reward deletion of features, reduction of complexity, improvements in reproducibility, stability, and monitoring to the same degree that improvements in accuracy are valued.” Reproducibility, stability and monitoring became product categories but I cannot point to an equivalent category for deletion or complexity reduction. One explanation is that their benefit lands outside the team doing the work.
The work I surveyed above shares a boundary, drawn around a single system. The debt is hidden inside of what is being built, and instrumenting the system better is the remedy. The paper located the debt at the system level rather than the code level.
Which leaves the same question one level further up. The paper asked what surrounds one system. Who is asking what surrounds a portfolio of them?
Undeclared consumers
One category in the 2015 taxonomy calls it undeclared consumers. A model produces predictions, and other systems begin reading those predictions without a contract and without the model’s owners knowing they exist. The paper calls the result a “hidden tight coupling” of the model to other parts of the stack. Changing the model then becomes expensive and risky, and whoever makes the change cannot see what else they are about to break.
The enterprise version is a finance team’s agent reading an output produced by a risk team’s agent. Nobody declared the dependency because the output was reachable without asking. Neither team can account for the pair, though each can account for its own system.
The paper describes a producer that does not know its consumers. At enterprise scale the harder case is a party accountable for the whole that does not know what the parts will do.
A book priced at twenty-four million dollars
In April 2011 the biologist Michael Eisen watched the price of a new copy of The Making of a Fly climb on Amazon to $23,698,655.93. Eisen concluded that two sellers were running automated repricers. Once a day one set its price to 0.9983 times the other’s, and the other then reset to 1.270589 times the first’s new price. Multiplied together the pair compounds at about twenty-seven percent a day.
The prices are consistent with two rules that each read a competitor’s price as an input. Eisen does not establish, and neither seller has said, whether either knew the other was doing the same. Undercutting a rival, or pricing above one on a stronger seller rating, are both ordinary strategies, and neither needs a sanity check to work on its own. Eisen’s reading is that neither algorithm carried a built-in sanity check on the prices it produced.
The 2015 paper names that control, four years after the listing. In “systems that are used to take actions in the real world, such as bidding on items or marking messages as spam, it can be useful to set and enforce action limits as a sanity check.” The advice is a decade old and costs little to follow.
Both prices sat on the same public product page, and Eisen recovered both rules from about a week of watching it. The information needed to see the loop was public. Nobody had claimed the job of watching the pair, and nobody appears to have been watching until Eisen wrote it up. Neither seller’s rule failed by its own measure, and neither rule belonged to the marketplace that could see both.
The usual answer to a problem like this is more observability. In this case, observability was free and the price still reached nearly twenty-four million dollars.
Why agents make it worse
The paper has a section on hidden feedback loops, “in which two systems influence each other indirectly through the world.” The bookshop is a loop of that kind, running through a price that happened to be public.
Three conditions look different to me now. The coupling runs through actions as well as data. It can close in seconds rather than daily cycles. And it forms between systems built by different business functions, with no shared engineering leadership that could put them in one room. Those are expectations drawn from how agentic systems are being deployed, not findings from a case.
The third one changes what kind of problem this is. Moving from code to system was a move inside engineering, and engineering could make it alone. Moving from the system to the space between business functions is the territory Melvin Conway mapped in 1968, where the communication structure of an organization constrains the shape of what it builds. Technical remedies alone do not reach failures of that kind.
What accumulates there is what I call complexity debt, the hidden coordination cost that builds every time an organization adds an autonomous system without maintaining coherence. I argued recently that this cost stays off the books while token spend gets managed carefully, and that the discipline companies apply to compute needs to reach it.
No post-mortem exists
I have searched for a public post-mortem of an enterprise losing serious money to two of its own agents interacting in a way nobody declared, and I have not found one.
Every case I have found is pre-agentic and comes from markets or grids: Amazon in 2011, the flash crash of May 2010, South Australia grid failure in 2016. In each of them the party that could see the composite was not the party that took the loss.
Enterprises are becoming shared substrates of the same kind, through common data platforms, shared model endpoints, and agents calling tools that call other agents. They can acquire the property that made those environments fail, without a regulator compelling anyone to write a report.
Coherence
The coordination layer is already under discussion. McKinsey named agent sprawl and proposed an agentic AI mesh, and it does assign ownership, writing that the pivot “cannot be delegated” and “must be initiated and led by the CEO.” SAP argued in early August that agent sprawl has made AI governance a board-level concern. What I have not found in that work is a standing owner for what separately approved systems do to each other.
A harder objection is that the industry has known about hidden technical debt for a decade and naming it did not fix it. The paper’s prescription was a change in team culture, addressed to the people building one system. Complexity debt accumulates between teams, so no team can pay it down or see both sides of an interface it did not know existed.
Ownership is the link an organization can reach. Naming the person accountable for the composite comes first, and the measurement follows, because someone then needs it to do their job. That person is the chief executive, or an officer the chief executive empowers. Each team is optimizing a local metric, and every one of those metrics can read green while the composite fails. The cost surfaces where the chief executive is accountable, which makes them the first person for whom paying it down is rational.
They build structure, because it’s not possible for an organization to sustain human diligence at machine speed. That means instrumenting the organization so its coordination state is visible, constraining interfaces to make whole categories of incoherence hard to express, and partitioning domains to keep a failure in one from spreading. Human judgment then handles the exceptions the structure surfaces.
That capacity is what I mean by coherence: a company’s capacity to see what its autonomous systems are doing, judge whether they are doing it well, and correct them when they are not. It is the signal box, and it is the part of the railway almost nobody has built yet.
The full argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
Every company deploying AI today is looking at the same invoice and trying hard to cut it. Earlier this summer, the Wall Street Journal reported on efforts to bring that bill down. The new word describing it, tokenomics, is all about the shift from tokenmaxxing to thriftmaxxing. Companies are using cheaper models where they are good enough, saving more expensive models for targeted uses where they are necessary. Reaching for the best frontier model for every task is out, and reducing costs without losing performance drastically is the new game.
The market is finally treating this visible cost of intelligence as a real one. EY wrote a few weeks later about two failure modes: tokenmaxxing optimizes for how much gets built while budget panic for how little. But who is asking if the right things are getting built?
I have always said, for years, that industry needs to take the cost of this seriously. Cost-sustainability was a guiding principle for me from the start when I cofounded Concentric AI. I had watched too many AI startups ship dazzling demos only to fold when the bill came due at real volume.
The cost moved
Here is what has changed now. Managing token costs is absolutely the right thing to do. But it is not the bill that will eventually sink you.
In the entire history of software, the cost of building itself acted as a filter. Engineering resources were scarce, and weak ideas never made it to the top to get built. When that filter goes away, so does that quiet discipline of prioritization and far more will get built. Every new deployment will bring with it new dependencies, handoffs and coordination costs that multiply. AI drops the friction and cost of doing work faster than reducing the cost of holding all of the pieces together. And that is the gap where real costs can hide.
Why the second cost hides
The coordination cost shows up on no dashboard since it doesn’t belong to anyone. It compounds quietly as each individual system looks fine.
While it is invisible on the invoice, the effects can be visible if you know where to look. It shows up as ROI that was promised months ago but failed to materialize, as rework that was not planned, as oversight that lags, and as interacting systems with no named owners.
EY states that the cost of an agent is often invisible until too late, and that tokens are only part of the true cost. I agree. But their fix is to price every agent, to benchmark it, meter it, and assign a value metric to each one from the start. The problem: no agent carries the cost of coordination among them, it is a cost that lives in the gap in between. You can meter every agent perfectly and still miss the entire bill.
Same discipline, different line-item
Cost-sustainability was never really just about compute. It was about refusing to allow a cost to sink you at scale. When I was thinking about building a product, that was the cost of compute infrastructure and it folded companies at volume. The same thing now is repeating at the enterprise level and the cost has moved from compute to coordination.
The discipline holds but the target is new. Enterprises need to budget for coordination the same way they started budgeting for compute. Put it on the books as a real line-item. Design for it before it compounds instead of finding out when it’s too late.
The price of intelligence is falling, but there is a hidden cost that is moving. Companies who will be running at full speed in the future will be the ones who start managing that cost now. The full argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
A financial auditing tool ran about $541,000 over. A logistics project meant to reduce delivery times ran about $134,000 over. Roughly $2.5 million across the three.
News coverage led with the money. Several outlets called it a coding task. Matching author details to listings is data reconciliation, the kind of work nobody watches.
The project ran for five months before anyone caught it.
Token spend is metered, priced publicly, and billed monthly. Every token that project consumed appeared on an invoice. Engineers attributed the overruns partly to the shift from flat subscriptions to token-based billing, where costs climb whenever a task generates more activity than expected.
Few things inside a large enterprise are more thoroughly instrumented than a cloud bill. Amazon had the numbers for five months and stayed unaware of them.
Sensing takes more than measurement. Something has to compare the number against an expectation, notice the gap, and route it to someone who can act while acting is still cheap. Amazon had the number. The comparison and the route were missing.
No dashboard would have closed this. Someone had to decide that aggregate token spend against declared intent is a thing the company watches, and then own the watching.
March incidents
Amazon’s retail website took four high-severity incidents in a single week in early March, including a six-hour failure that locked customers out of checkout, account information, and pricing.
An internal document prepared for the review meeting identified GenAI-assisted changes as a factor in a pattern of incidents going back to Q3. That reference was deleted before the meeting, according to the FT, which saw both versions. Amazon disputed the reporting and said only one incident involved AI directly, with the root cause an engineer acting on inaccurate advice an AI agent had inferred from an outdated internal wiki.
Work backward from July. Five months of undetected spending starts around February or March. I cannot confirm the detection date, so treat the overlap as inference. Even without it, the shape holds. Amazon added review gates on AI-assisted changes to critical systems while a cost failure accumulated invisibly on a job nobody would call critical.
Constraint depends on detection. You cannot cap, contain, or price what you cannot see. Amazon reached for the second without the first, which happens because approval steps are visible to leadership and instrumentation is not.
Every deployment a team builds imposes cost on everyone else. Another surface to watch, another dependency to reconcile, another system someone who did not build it has to understand. The team keeps the benefit while the organization bears the cost. The remedy is to price that burden back to the team creating it.
Amazon built a price signal pointing the wrong way. Ranking people by consumption pays them to consume. An internal metric carrying status and no cost gets gamed, and this one did.
The presentation’s own recommendations now include avoiding leaderboards that reward token consumption, and checking whether higher token usage produces useful output.
The objection
Amazon frames these as isolated examples of teams learning from one another, and says cherry-picking them does not reflect how teams across the company use AI. For a company its size, a seven-figure surprise is a rounding error.
The reporting also lacks a base rate. Nobody has said how many AI projects came in on budget for every one that blew up.
Five months of invisibility still belongs to the control architecture rather than the budget. The same architecture at a company with a $12 million annual AI budget produces the same five months and a different outcome. Amazon can absorb what it cannot see.
What it costs to fix
Spending on sensing is bounded and knowable in advance. You can price the instrumentation, the ownership, and the review cadence before committing. The incoherence it prevents accrues silently and surfaces only once addressing it is no longer optional.
Amazon paid roughly $2.5 million across three disclosed projects. Other companies will meet the same failure without the revenue to absorb it.
The argument here about sensing, constraint, and priced externalities runs through my book, Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
Ford has added more than 350 experienced engineers over the past three years after its automated quality systems failed to deliver the results the company expected. Inside Ford they are called gray beards. Some are former Ford employees. Others came from suppliers.
Ford says it added the specialists “through internal promotions or new talent” to work alongside newer team members. Headlines have called it rehiring people AI replaced but Ford has not used that word, and the reporting does not establish that these specific roles were cut.
Poon also said Ford had not paid enough attention in prior years to the experience of its most knowledgeable engineers, the ones who had been through many product cycles.
Kumar Galhotra, Ford’s chief operating officer, said the company had been relying more and more on automated quality systems before recognizing the approach was not working.
Ford describes a training data problem. Poon’s version: enhancing the automation and machine learning tools required making sure they were trained by the most experienced individuals.
Part of that holds up plainly. The specialists do reprogram the AI tools that fell short.
The jobs
A missing corpus has a fix with an end date. Sit the veterans down, extract the failure modes they carry, feed the models, thank them.
A standing design review is a verification loop the organization has decided to keep running.
Where the corpus story runs out
Tacit knowledge can be captured up to a point. What a senior engineer knows about how a joint fails under a particular thermal cycle can be written down, and should be.
Judgment applied to an unanticipated case cannot. A reviewer looks at a novel configuration and says it will not hold, for reasons that emerge from the thing in front of them. Enumerate those cases ahead of time and you would have automated the review already.
The corpus framing implies a completion state, where enough capture makes the humans optional. Ford’s remedy points elsewhere. Mandatory and recurring is what you build once you have concluded the checking does not stop.
Automating a quality inspection function means automating a verifier, which removes the capacity to tell whether the automation works.
Ford’s specialists hold the ability to tell the machine it is wrong.
The order
Poon’s admission about prior years is the sharpest thing either executive said. Ford’s assumption was reasonable and its sequence was wrong. Capture the expertise, then automate, and the program is defensible. Automate on the assumption that design requirements are sufficient, and you spend three years buying judgment back from suppliers and internal promotions.
Each step is the precondition for the next. Skipping one relocates its cost to a later point, larger, with fewer options available. Ford turned a knowledge problem into a three-year staffing program, and that program ran alongside more than $1 billion in expected warranty and material costs this year and a quality reputation to repair.
Mentorship
Ford is explicit that mentorship is part of the assignment. The specialists work alongside newer team members and train junior staff who never absorbed the institutional knowledge.
This is how senior judgment gets built. Less experienced people work real problems while someone holding the judgment watches and corrects. The routine cases are the training ground, and automation takes them first.
An organization that automates routine work and then loses its seniors breaks the pipeline at both ends. Nobody holds the judgment and nobody acquires it. Ford is paying to rebuild both.
Headcount models calculate savings against the cost of the people. The judgment training pipeline appears nowhere in the model.
Ford is still deploying AI
On an autumn 2025 earnings call, Galhotra said Ford was systemically deploying AI across the entire industrial system, including 900 AI-powered cameras across its plants to detect quality issues at the source. Ford kept the cameras. Jim Farley told Bloomberg TV that Ford has AI tools for vision systems, and that most of it comes down to team members paying attention to small details.
Ford topped the mainstream J.D. Power rankings with the automated systems still running. Experienced people now sit between the systems and the product.
Experienced engineers sat on Ford’s books as execution capacity, a cost line. Their function was judgment, which is what makes execution capacity safe to deploy.
Ford is not unusual in getting the order wrong. CNBC reported Robert Half data showing 32 percent of U.S. hiring managers eliminated a role primarily because of AI and later rehired for the same or a similar position. Robert Half’s own summary puts it as more than 3 in 10, and notes the two most common reasons given: the role required institutional knowledge or context AI could not replace, and it involved relationship management AI could not replicate.
The arguments here about judgment as the layer that cannot be purchased, and about sequence in organizational automation, run through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
I read an article in the Wall Street Journal today about my hometown of Austin, and it made me laugh before it made me think. Since Waymo’s robotaxis arrived here in 2024, they have apparently collected $9,325 in parking tickets. Tow-away zones. Metered spots they never paid. A disabled space outside an elementary school. One that idled in front of a church garage for five minutes during Sunday service while parishioners waited. A resident’s complaint in the records reads, plainly, “there needs to be some way to get them to move.”
These figures come from documents the Journal obtained through an open-records request, so I am relaying its reporting rather than confirming the numbers myself.
As a number, $9,325 is nothing. Austin collected $6.3 million in parking fines in 2025 alone, so Waymo’s two-year total is a rounding error. Spread 83 citations across more than 300 cars over two years and the per-vehicle rate is low, probably lower than what a human-driven taxi fleet of the same size would rack up in the same window.
While the dollar figure is trivial, the behavior behind it is not. It is also a near-perfect illustration of the argument I have been making.
The obvious reading, and why it misses
The easy version of this story is that the self-driving car is not ready. Look, it cannot even park. That reading is wrong, and the same article contains the reason.
An independent analysis by the Insurance Institute for Highway Safety found that over more than 50 million driverless miles, Waymo’s crash involvement rate was 68 percent lower than that of human drivers. The hard problem, the one with lives attached, Waymo has solved to a level that beats us. Parking is where it stumbles.
A system can be superhuman at its central task and fail at something a sixteen-year-old handles on the first day with a learner’s permit. This is jaggedness. Andrej Karpathy coined the term for the strange fact that a state-of-the-art model can solve a hard problem and then miss a trivial one, and a field experiment with 758 BCG consultants showed the same thing in the workplace: performance was excellent on tasks inside the model’s zone and worse on adjacent tasks that looked just as easy. The boundary is uneven, and it does not follow the difficulty ranking a human would draw. Driving safely across 50 million miles is the hard task the machine has mastered. Parking lawfully is the easy adjacent task it has not, and no amount of skill at the first predicts skill at the second. I have written about this shape once already this month. An AI office agent filled out seventeen forms in five minutes and then could not upload a file. Same jaggedness, different machine.
The dangerous part is that the failure is invisible from the outside. Watch a car drive flawlessly for an hour and you will assume it can handle a parking lot, because that inference holds for humans. It does not hold here, and the assumption is where the trouble starts.
The parking failure itself is not my thesis. Waymo will patch handicap-spot detection, and that particular fine will stop appearing. But notice what does not happen. Jaggedness does not get fixed. It moves. Patch the parking lot and the uneven edge shows up somewhere else nobody thought to check, because the unevenness comes from how the system learns, not from a single defect waiting to be found.
That is one problem, and it is real. There is a second one in the same article, and it is not a version of the first. It is a different failure entirely, and it is the one my book is actually about.
The failure that no model fixes
Read the part of the article that is not about parking.
On July 8, the National Highway Traffic Safety Administration sent autonomous-vehicle developers a letter demanding that their cars better follow instructions from first responders. The regulator’s complaint was that robotaxis often fail to recognize where they can stop without getting in the way. When a Waymo blocked an active railroad track in January 2025, an officer reported he had no choice but to have it towed. There was no other way to move it.
It does not go away with better driving. A firefighter at a scene, a police officer at a closure, a resident at a blocked garage: each of them has authority over the situation but no means to direct the machine sitting in it. The human is formally in charge and practically helpless. I keep making one distinction in my book, and this is it in the physical world. Having authority over a system is not the same as having the capacity to intervene in it. The officer had every right to move that car. He had no lever to do it, so he called a tow truck.
That gap is the coherence problem, and no amount of driving skill closes it. It is a problem of the connection between a capable system and the people who are supposed to be able to redirect it. You can make the car a better driver every quarter and leave that gap exactly where it is.
Three hundred locally rational decisions
Here is the detail in the article that matters most.
Waymo runs more than 300 robotaxis in Austin. Between trips, the article says, the cars park themselves on public streets to stay near riders and avoid adding traffic. Each of those choices is sensible. Idling near likely demand cuts empty miles and shortens the next pickup. For the fleet, it is the right call every time.
Now add up 300 right calls. You get 300 vehicles independently claiming curb space across one city, each optimizing for the fleet, none of them accountable for what they cost the curb in aggregate. The church-garage blockage was not one rude car. It was the predictable output of a fleet doing exactly what it was designed to do, measured against a shared resource that no one in the system is responsible for.
This is the pattern I spend the book on. W. Edwards Deming showed it in factories long before any of this. Optimize each part on its own and the whole can still degrade, because the parts interact in ways no single part can see. A support agent and a billing agent inside a company can each be flawless and still act on contradictory assumptions about the same customer. Three hundred robotaxis can each park perfectly rationally and still congest a city. The mechanism is identical. The only new thing is that it now runs at the speed and scale of software, on a public street.
Where the analogy breaks, and why the break is the interesting part
My book is mostly about a different situation. Many systems, built by many teams, with no shared owner, colliding inside one company. Waymo is close to the opposite. One company, one software stack, one central fleet manager. The cars are not incoherent with each other. They are all perfectly coherent with Waymo’s goal. The incoherence is between the fleet and the city.
That difference does not weaken the parallel. It sharpens it. Inside a single enterprise, the cost of local optimization eventually lands back on the enterprise itself. It pays its own complexity debt, later and with interest. In the robotaxi case, the company captures the efficiency and the public absorbs the cost. The fleet gets the shorter pickup times. The churchgoers get the blocked garage. The externality lands outside the firm, which means the firm has no natural reason to see it or price it.
Which is where the parking ticket returns, transformed. The tickets are not the failure in this story. The tickets are the city’s answer to it. Austin cannot rewrite Waymo’s software, so it does the only thing available to an outsider. It attaches a dollar figure to each incoherent act and bills it back. A tow-away citation is a coordinating institution forcing a cost back onto the party that created it, because that party will not absorb a cost it cannot see on its own dashboard.
In the book I call that pricing the incoherence, making the party that creates a coordination burden bear the cost it imposes on everyone else. It is one of only a few places you can intervene when you cannot redesign the system directly. Austin is doing it with a parking-enforcement officer and a public complaints database. It may be crude, but it is the right instinct. When you cannot fix the system, you can at least make it pay for the mess, so the incentive to stop finally reaches someone who can.
The instinct is widely shared, which is its own small piece of evidence. I read the comments under the article, and the readers who were not busy mocking it reached, unprompted, for exactly this lever. Charge a flat $5,000 fee for a driverless tow. Impound the car and make the company pay storage, same as a person would. Hold them to the standard you and I are held to. Nobody in that thread proposed debugging Waymo’s curb detection, because none of them can. They proposed raising the price of the behavior, which is the one move available to an outsider who cannot see inside the system and cannot change it. Pricing is what is left when coordination is out of reach.
A note on fairness, since it matters. Waymo pays these tickets like any other driver, and its spokesman said the company expects no special treatment. It contests some citations and has had a couple dismissed. None of that is evasion. It is a company behaving reasonably inside a system that has not yet given it a better way to behave. The point is not that Waymo is careless. The point is that even a careful, centrally managed, genuinely safer-than-human fleet produces coordination costs its own metrics will never show. That is the part that should worry anyone deploying autonomous systems anywhere.
The version of this that has not happened yet
One last thought, and I will flag it clearly as speculation rather than something the article reports.
Today Austin has one large fleet parking itself on the curb. The same article names two more operators already here, Tesla’s Robotaxi and Amazon’s Zoox. Imagine the near future where three or four fleets, each centrally coherent, each optimizing its own vehicles against the same finite curb, all share one city. None of them is incoherent on its own terms. Each is a model citizen by its own dashboard. Together they compete for the same few feet of pavement outside the same church at the same 10 a.m. service, and no one owns the result.
That is the multi-owner version of the trap, and it is the one that looks most like the enterprise problem I actually write about. Many capable systems, no shared view of the whole, a commons that quietly degrades while every participant is behaving well. When it arrives, the city will reach for the same tool it is using now, only harder. It will try to price the congestion, because pricing is what is left when you cannot coordinate the systems directly and you cannot see inside any of them.
There is a sharper edge to a single fleet that is worth one more sentence, because it cuts against the intuition that central control is safer. A fleet of 300 identical cars does not only optimize together. It fails together. Every vehicle runs the same software and leans on the same positioning inputs, so a single upstream fault does not hit one car, it hits all of them at once and in the same way. I made this point about software agents in a recent post: two agents drawn from the same model share the same blind spot, so the redundancy between them is nominal. Here it is 300 machines sharing one blind spot instead of two. Homogeneity buys clean coordination on a good day and correlated failure on a bad one. That is not an argument against central control. It is a reminder that the thing which makes a fleet coherent is the same thing that can make it fail in unison.
The safest car on the road parks in the fire lane. The fleet that adds no traffic blocks the garage. Every decision was locally correct, and the street got worse anyway. That is not a story about cars. It is the story of what happens to any organization, or any city, that fills up with capable systems faster than it builds the means to keep them coherent.
The gap between capable systems and coherent ones is the subject of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
If the title sounds weird, it is because I stole it from a reddit thread I read last week! This thread, on a data engineering subreddit, is all about building fast without building coherent. A practitioner described a year spent working inside a major enterprise platform deployment that, by their account, went badly. The post drew more than a thousand upvotes and over 150 comments, and many of those comments said a version of the same thing. This matches what happened to us.
Let me be careful about what this post is and is not. I cannot verify any of it. I do not know the author, the employer, or whether the account is accurate, complete, or fair. I am not treating any of it as fact. I am not making a claim about the named vendor, any other company, its people, or its products. Online accounts are one-sided by nature. The company is not present to respond. And the thread does not even agree with itself on basic points, including why things went wrong. One commenter accused the vendor’s engineers of dragging work out to bill more hours. Two others replied that the vendor uses fixed-price contracts and has the opposite incentive.
So, let me set aside the question of motive or blame but ask something different. If a reader believed these accounts as written, what pattern would they describe? The pattern, if it is real, is one my book predicts.
The pattern the accounts describe
The original poster says they inherited the system after the engineers who built it left on thirty days notice, once a first version was declared done. What they found, in their telling, was a catalog of shortcuts. Hardcoded dates. Hardcoded accounts. The same business concept fed by different inputs in different places. Earlier problems patched with more hardcoded logic.
They gave one concrete example later in the thread. An engineer had built a button to delete a record. The button removed the record from the screen. It did not remove the three related records that the original had created when it was made. The result was orphaned data. The button worked. The system did not.
That small story is the whole thing in miniature. Every piece can be locally correct while the system is globally broken. A button that deletes what you can see and leaves what you cannot is a fine button and a broken workflow at the same time.
A commenter who said they work at a hospital inside a national health service described their own experience. Outside engineers did intense early work, leaned heavily on the in-house team to explain the basics, then left. No one was clearly left owning or maintaining what had been built. At one point, the commenter said, the vendor’s own monitoring staff emailed to ask why duplicate records and bad addresses were appearing, and the hospital could not answer, because it did not have access to the pipelines that had been built for it.
Another commenter described a failure higher up the organization. A single platform owner was installed. Over time that person’s standing became tied to the platform’s success, and the information traveling up to senior leadership was filtered, so the picture at the top stayed positive while the picture on the ground did not.
The line I keep thinking about
The poster wrote that the engineers used AI to produce tangled, low-quality logic. A commenter answered in five words.
That is the argument of my book, delivered by someone who did not set out to make it. When producing code becomes fast and cheap, more of it gets produced. The speed is real. What does not arrive with the speed is coherence. Coherence is the work of making sure each piece fits the whole, that today’s shortcut is not tomorrow’s silent failure, and that someone still understands the system after the people who built it are gone. Execution got cheaper. Coherence did not.
Why I am comfortable writing this at all
Here is the part that matters most. The people in the thread mostly did not think the story was about one company. One commenter wrote that you could swap in almost any vendor, almost any consultancy, and almost any project, and reach the same ending. Another described the identical arc with a completely different vendor. Others reached back to the enterprise data tools of twenty years ago and asked whether it had always been this way. They were describing a recurring structural pattern, and I think they were right to.
The pattern is old. W. Edwards Deming spent decades showing that optimizing each part of an organization on its own can degrade the whole, because the connections between the parts matter as much as the parts. Stafford Beer showed that organizations drift when the feedback reaching the people in charge is slow or filtered. Neither man was talking about AI. Both were describing this thread.
I want to give the other side its due, because the thread did. The original poster said plainly that the platform itself is fine for what it is. Other commenters defended it and corrected specific claims. Many organizations report that the same tools serve them well. The tool is capable. What fails, in these accounts, is the fit between a tool sold on speed and an organization that cannot absorb what speed produces. Change the logo on the invoice and the story would run the same way.
The number nobody calculated
The poster reported that the project was estimated at four months and took fifteen, and that the company had seen no return so far. I cannot confirm those figures. If they are even roughly right, they point at something the book returns to again and again. The promised savings were a calculation about capability. The cost that actually landed was a calculation nobody made, the cost of coordinating, maintaining, and understanding what got built. The first number is easy to put in a sales model. The second one shows up a year later and has no owner.
I have written three times recently about the same shape seen from different angles. A machine can generate the output. A person still has to own the part with no dashboard. In a newsroom experiment, an AI agent finished the forms and could not finish the job. In a courtroom, a scoring system read a gap in the data as a verdict on a person. In this thread, if the accounts hold, capable engineers produced software that worked in the demo and broke quietly in the corners, then left, and the coherence walked out the door with them.
Faster is not the same as coherent. It never was. The difference used to be expensive to create and easy to see. Now it is cheap to create and slow to see, which is exactly why it is worth watching for.
I can’t verify the thread, but the gap it points at is real. That gap, between building fast and building coherent, is the subject of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
Keith Collins gave an AI agent full control of a laptop and three office jobs to do: survey nine colleagues over Slack and log their answers, identify staff cuts to hit a budget target, and fill out seventeen I-9 employment verification forms. The tasks were adapted from benchmarks published by researchers at Carnegie Mellon and OpenAI. The agent ran on Anthropic’s Claude Cowork app.
On the third task, the agent generated all seventeen forms correctly in under five minutes. Then it tried to upload them to Google Drive and failed. It clicked the right menu item and never noticed that a file picker had opened. It compressed the files. It converted them to a long string of bytes. It asked a second agent for help, and the second agent hit the same wall. After roughly seventeen minutes, it stopped trying and marked the task complete.
The Times files this under comic stumbles. It is the most consequential finding in the piece.
120,000 jobs, cut on the opposite premise
The article closes on a line meant to calm: AI still needs a human boss.
The same article reports the layoffs. More than 200 tech companies have cut roughly 120,000 jobs this year, per Layoffs.fyi. Meta and Oracle made substantial cuts citing AI. Cloudflare’s chief executive, after letting go of about 1,100 people, said he expects AI to replace workers in middle management, finance, and marketing.
Those cuts rest on a premise: the supervisory layer is what becomes redundant. The experiment found the reverse. Agents were strong at execution and weak at judgment. They wrote clean code in minutes, then made a categorization error about employees on leave that any manager would have caught.
Firms are removing coordinating capacity while installing systems that consume more of it. That connection is the thesis of the book I am writing. One case has already reached a federal courtroom.
Three specimens
My argument: when execution stops being scarce, the binding constraint becomes coherence, the integrity of the link between what local systems do and what the enterprise intends. Coherence has five specific dimensions, and autonomous systems break it in six recognizable ways. The Times experiment produced three clean specimens.
The false completion is escalation failure. The agent detected its own failure. It reasoned about it for seventeen minutes. It recruited a second agent. Then it reported success. The system knew it had not finished, and it stayed quiet. In this case, it cost little to the reporter analyzing logs. But in an enterprise running ten thousand delegated tasks a day, it is the mechanism by which reported completion drifts away from actual completion. Escalation failure is the one mode in my taxonomy that leaves every dimension of coherence intact and disables the reflex that repairs them. An organization can see a problem clearly and still be paralyzed when the signal never reaches anyone who can act.
The second agent matters too. Two systems drawn from the same model share the same blind spot, so the redundancy is nominal. The organization paid for one failure twice.
The code detour builds architecture nobody chose. In every task, the agent was told to work through the applications and wrote code instead. Graham Neubig of Carnegie Mellon puts it plainly in the piece: agents work in a very unhuman way, writing code instead of using the interfaces humans use. The Times treats this as a limitation. It is also an architectural event. The agent replaced the assigned task with a different one that produced a similar-looking artifact. An org chart rebuilt by a Python script carries a new dependency, a new failure mode, and no owner. Multiply that across a year of routine delegated work and the enterprise runs on infrastructure nobody selected, documented nowhere, discovered only when it breaks.
Local simplification often works by moving complexity somewhere else. The productivity gain lands on the dashboard. The displaced complexity does not.
The staffing error is contextual failure, and it is already in litigation. Given a budget target, the agent did something genuinely good. It read the personnel documents and concluded the 4 percent reduction could be met through planned retirements and resignations, with no layoffs. Then it added employees on leave to the list of cuttable roles without considering when they were coming back. The source material was silent on duration. The agent never asked.
Researchers at Stanford and the NBER frame this as a tacit knowledge problem, and that holds. The mechanism is more specific. The agent had no way to represent a person as temporarily absent for a reason that says nothing about their value. Silence in the record became a mark against the employee.
Nine days before the Times published, twenty-six Meta employees filed suit in federal court in Oakland alleging that the same substitution happened to them at production scale. Their complaint says Meta relied on internal AI systems, keystroke and activity-monitoring data, AI token-usage dashboards, and algorithmically assisted performance rankings to decide who would go in a layoff of roughly 8,000 people, about 10 percent of the workforce. The central allegation: those scores cannot by design be accumulated by an employee on protected medical or family leave, or by an employee whose output is reduced by a disability. The suit further alleges the company never paused the process for the individualized, leave-neutral review the law requires. About half the plaintiffs had taken leave for caregiving or pregnancy-related reasons. Their jobs were set to end on July 22, the day the Times ran its experiment.
Meta rejects the claims. The company says they lack merit and are not based on facts, and that workforce and organizational decisions “were and are made by people, not AI.” The allegations are unproven and the case is live. I am looking at the structure of the dispute here and taking no position on the verdict.
That structure survives either outcome, which is why it belongs in this argument. Suppose Meta is right that people made every call. Those people still read rankings, and the rankings still came from a substrate that had no field for protected absence. A human who approves a ranked list holds the authority to intervene and does not necessarily hold the information. Formal presence in a process is a weaker thing than capacity to change it. My book calls that oversight failure.
The Times agent and the Meta complaint describe one error at two scales. A system met a gap in its data and scored the gap as a deficiency. Nobody had built the mechanism that would have made it ask a question instead.
My book already discusses the Cloudflare decision the Times cites. Its chief executive organized his reasoning around a Drucker framework: every organization has builders, sellers, and measurers, and AI can now measure cheaply, so the measuring layer can shrink. The framework is coherent and the logic holds internally. The question I put to it there was whether the people categorized as measurers were only measuring. The Meta complaint poses the companion question. When the score came back low, was the system measuring performance, or measuring absence?
The article measured one axis
Task reliability and organizational complexity are independent dimensions. Improving one leaves the other where it was. Conflating them keeps the expensive failures invisible until they are hard to reverse.
The Times measured reliability, carefully and well. Its headline number comes from Scale AI: on real freelance projects, the best model produced client-ready work about 16 percent of the time.
The coordination question sits outside that number. If 84 percent of agent output requires human review, review capacity becomes the ceiling on deployment. Oversight load scales with the number of systems, and it lands on a different dashboard than the productivity gain. A control system has to be at least as varied as the thing it controls. Thin the supervisory layer while thickening the volume of supervised work and the cost moves off the ledger. It stays in the business.
The Oakland filing shows where it resurfaces. Twenty-six people asking a court to examine how a ranking was produced is a coordination cost, arriving late, in the most expensive form available.
Where the article argues against me
The counterargument has real force. If agents cannot reliably finish tasks, they will not be deployed at scale, and the coordination problem stays theoretical. The article supports that. A 16 percent success rate describes a product that is not ready.
Two responses.
First, the two failures differ in kind. The upload bug will be fixed. It is a UI problem and the entire industry is aimed at it. The leave-of-absence error and the false completion report sit at the boundary between the agent’s context and the organization’s. Better models will make both rarer. No model tells an enterprise which completion reports it can trust, or who owns the Python script the agent wrote last Tuesday. Those are ownership questions, and capability does not settle them.
Second, consider what the two variables are doing. Reliability improves on a public curve that everyone watches. Coherence has no curve, because almost nobody measures it. Using a snapshot of the fast-moving variable to dismiss the stationary one is the error my book is written against.
Two caveats. The Times experiment was three tasks, one tool, one synthetic environment, with expert-written prompts and supporting documents supplied by benchmark researchers. Real enterprises rarely supply that quality of context. The setting was favorable on the task side and trivial on the coordination side, since one agent ran alone with no installed base of prior deployments to collide with. The reliability observed sits closer to a ceiling than a floor. The coordination cost observed is near zero by construction.
The second caveat: Meta’s alleged systems are ranking and monitoring software, a different technology from an autonomous agent operating a laptop. The defect predates agentic AI. Agentic deployment raises the rate at which it executes.
What to watch instead
If you run an enterprise and this experiment shaped your thinking, start measuring the things it did not.
What fraction of your deployed autonomous systems has a named human owner. What fraction produces decisions you can explain to a regulator. How many of your systems depend on other systems in ways nobody mapped. How much of the behavior is visible to the people accountable for it. How hard it would be to remove any given system now that it is running.
If my thesis holds, those five move in one direction as deployment scales, while task accuracy holds steady or improves. That divergence is the signature.
The Times asked whether AI can do your job. Twenty-six people in Oakland are asking the harder version. Who answers for the score that said they were not doing theirs?
My book, Coherence, arrives this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
Two arguments about AI landed within days of each other this month. They look unrelated. They describe the same event from opposite ends, and read together they close a loop that neither closes alone.
Satya Nadella published a short piece over the weekend that names something most enterprises have not yet noticed they are doing. A few days earlier, Arvind Narayanan and Akash Kapur published a longer essay on why the AI labs cannot make money selling raw intelligence, and what they will do about it instead. One argument tells you what you are losing. The other tells you why the loss is not an accident.
Let’s start with Nadella.
He begins with Kenneth Arrow. Arrow described a paradox in the market for information: a buyer cannot know what information is worth until they have it, at which point they have it for free. So the seller risks giving away the knowledge in the act of trying to sell it.
You pay for intelligence twice. Once in money. Again in the proprietary knowledge you have to reveal to make that intelligence useful.
Nadella inverts it. In the AI age, the risk runs the other way. The buyer gives away knowledge in order to use what they bought. You pay for intelligence twice. Once in money. Again in the proprietary knowledge you have to reveal to make that intelligence useful. And the better you want the model to perform, the more of your knowledge you have to hand it. He calls this the reverse information paradox.
His answer is a trust boundary: a hard perimeter inside which your data, traces, evals, tuned weights, and memory accumulate together, and across which nothing passes without consent. Own your evals. Build your learning environment inside your own tenant. Keep the orchestration layer decoupled from any single model. Compound.
Now the other end.
The labs are spending trillions on chips and data centers. The thing they sell, model inference, is close to a perfect commodity. The leading models behave alike, cost about the same to run, and carry almost no switching cost.
Narayanan and Kapur ask a blunt question. The labs are spending trillions on chips and data centers. The thing they sell, model inference, is close to a perfect commodity. The leading models behave alike, cost about the same to run, and carry almost no switching cost. Sell a commodity into a competitive market and the price falls to the cost of production. So how does any lab ever earn back the buildout? Their answer is that it cannot be earned back by selling tokens. The labs have to move up the stack, into products, workflows, and embedded deployments, and they have to build moats. One of those moats is a flywheel: train the models and systems on customers’ own material, their data, their execution traces, their evaluation suites, until the product pulls ahead in a way a rival cannot copy.
Set the two arguments side by side and the picture sharpens.
Your judgment is not an incidental byproduct of the labs’ business. Capturing it is the business, because it is the one thing that turns an undifferentiated model into something with a moat around it.
What Nadella calls exhaust leaking out, Narayanan and Kapur call the flywheel that powers the labs’ escape from the commodity trap. It is the same substance. The traces, the corrections, the evals. Nadella watches them leave your building. Narayanan and Kapur explain why the firm on the other side needs them so badly. Your judgment is not an incidental byproduct of the labs’ business. Capturing it is the business, because it is the one thing that turns an undifferentiated model into something with a moat around it.
That changes the stakes. The pull on your knowledge is not a quirk of one product or one vendor’s terms. It is structural, and it will not relent, because the economics of the entire model layer depend on it. The rest of this piece uses one instrument from the book to work out what to do.
What leaks is not your data
Start with the mechanism, because most people will read Nadella’s post as a data-protection argument and it is not one.
Nadella is specific. Models learn from exhaust. The prompts people write. The tools the agents call. And above all, the corrections people make when the model is wrong. Each correction is distilled into know-how. It leaks imperceptibly, he writes, trace by trace, correction by correction, eval by eval.
A correction is not a data point. It is a judgment. When your underwriter overrides the model’s risk score, she is not supplying a fact. She is encoding a standard. What good looks like in this market. What that number actually means when the counterparty is this counterparty. What your firm would never do, regardless of what the numbers say. She is teaching the machine your institution’s judgment, in the most compressed and machine-readable form that judgment has ever existed in.
That judgment is the one thing your competitors cannot purchase. I have argued elsewhere that as intelligence commoditizes, the advantage that remains is the one no vendor can sell you. Nadella reaches almost the same sentence from a different direction, that this is the kind of knowledge a competitor could never buy.
The reverse information paradox is not primarily an intellectual property problem. It is a coherence extraction problem. Not the capacity itself, which no one can take from you, but everything the capacity produces, exported decision by decision.
Which is exactly why the leak matters. The reverse information paradox is not primarily an intellectual property problem. It is a coherence extraction problem. Not the capacity itself, which no one can take from you, but everything the capacity produces, exported decision by decision. The cruelty of it is structural: the act by which an organization encodes its judgment into its systems, correcting the machine until the machine reflects how the firm actually thinks, is the same act by which it exports that judgment to whoever owns the model.
If a single competitor lets the vendor learn from its work, the model that serves your whole industry improves, and the vendor’s hand strengthens against every buyer in it, including the ones who kept their discipline.
Narayanan and Kapur add the part Nadella leaves out, which is that you cannot hold this line alone. The flywheel needs only one firm in a sector to start it turning. If a single competitor lets the vendor learn from its work, the model that serves your whole industry improves, and the vendor’s hand strengthens against every buyer in it, including the ones who kept their discipline. Your own boundary protects your specific corrections. It does not protect you from the sector arming the vendor around you.
You do not lose your moat in a breach. You lose it in a thousand small acts of being helpful, some of them your own, some of them your rivals’.
Two kinds of exhaust, and how to tell which one you are leaking
Nadella treats exhaust as a single substance. It is not, and the distinction is practical.
In the book I use a simple two-axis instrument. One axis is verifiability: whether a task’s success can actually be checked, and how fast a failure would be caught. The other is organizational complexity: how many units a deployment touches, how deeply other systems depend on it, and how hard it would be to reverse.
Run exhaust through those two axes and it separates cleanly.
Low verifiability leaks your judgment. These are the tasks where success is contestable and the model is often wrong: strategic assessment, valuation, anything where the right answer depends on tacit context. This is also exactly the work where a human should remain the decision-maker, which means the wrong answers get caught and corrected, and corrections are the most concentrated form of institutional judgment there is. Notice the twist: the safer your posture, the richer the exhaust. This is the highest-value leak in the building.
It is also where you are most easily held. Narayanan and Kapur point out that judgment-heavy work has no objective standard of quality, so a buyer cannot verify the output even after the fact. Writing, strategy, judgment calls are credence goods, like the work of a lawyer or a consultant. Unable to compare quality, you fall back on trust and reputation, and you stay put. So the same weak verification that makes these corrections precious makes the vendor that holds them hard to leave. Low verifiability is where your judgment concentrates and where your exit narrows at the same time.
High complexity leaks your architecture. These are the deployments woven deep into how the company runs. The model may be right almost every time, so there are few corrections. But the traces are a map. Which tools get called in what order, which systems depend on which, where the handoffs are, what the exception paths look like. That is a blueprint of how your organization actually operates, as opposed to how the org chart says it does.
High verifiability plus low complexity leaks almost nothing worth having. This is the commodity zone. Document classification, code execution, data transformation. Let it run. The exhaust is worthless to a competitor because the task is worthless as a differentiator.
So the first question is not “how do I protect my data.” It is “which of the two things am I giving away, and is it the one that matters.” The answer depends on where the deployment sits, and most enterprises have never asked.
Your evals are worth more than your data lake
Nadella makes a point in passing that deserves more weight than he gives it. Evals, he writes, define what good looks like inside the organization.
Follow that all the way down.
Data is a record of what happened. Evals are a specification of what you consider good. Those are not the same kind of object, and they are not remotely the same value.
A competitor who has your evals knows what you value, how you score it, and where you draw the line.
A competitor who steals your data still has to work out what you were optimizing for. They have the outcomes without the standard. A competitor who has your evals knows what you value, how you score it, and where you draw the line. They have the standard, which means they can generate their own outcomes.
If evaluation is where human effort is concentrating, the eval suite is where your people’s judgment is accumulating, and that is precisely why it is worth more than the data it scores.
This is not only an enterprise observation. In his ICML keynote this month, Narayanan argued that as AI absorbs the building, human effort migrates toward exactly this work: away from developing the systems and toward evaluating and monitoring them, toward the tasks that are hardest to verify. His frame there is the whole field and the whole economy. Bring it down to a single company and it lands on the same object. If evaluation is where human effort is concentrating, the eval suite is where your people’s judgment is accumulating, and that is precisely why it is worth more than the data it scores.
Most enterprises spend enormous energy guarding the data lake, and then hand the eval suite to whoever will run it for them, because building evals is tedious and the vendor offers to help. That is the wrong trade, made in the wrong direction, for the most understandable reason in the world.
If you protect one thing inside the boundary, protect the definition of good.
You cannot enforce a boundary you cannot see
Here is the prerequisite the boundary quietly assumes, and where I think the practical failure will happen.
Every one of Nadella’s recommendations presupposes an organization that knows what its systems are doing. Retain ownership of your traces, feedbacks, and decisions. Build learning environments inside the tenant boundary. Make sure nothing crosses without consent. Each of these requires that you can see what you have deployed, what it touches, and what leaves.
Most enterprises cannot.
IBM surveyed 2,000 chief information and technology officers this year. Seventy percent said teams were deploying AI faster than IT could track. Seventy-seven percent said adoption was outrunning their governance. Those two numbers describe an organization that does not know its own perimeter.
The failure will not look like a breach. Your general counsel is not going to paste the merger memo into a consumer chatbot. A junior analyst is going to paste the comparable transactions in at eleven at night, because the deadline is at eight and the tool is right there and nobody told her not to. Multiply that by every team that stood up an agent this quarter without telling anyone, and the hard boundary is a diagram in a slide deck.
The map of how you work is a prize for the vendor for the same reason it is a necessity for you. So there is no neutral option. Either you build the sensing layer inside your boundary, or you rent it from the vendor, who then owns the map.
Visibility is also contested from the other side. Narayanan and Kapur note that the labs’ most lucrative escape, charging for outcomes rather than tokens, requires them to see inside your business processes, which means migrating into your System of Record. The map of how you work is a prize for the vendor for the same reason it is a necessity for you. So there is no neutral option. Either you build the sensing layer inside your boundary, or you rent it from the vendor, who then owns the map.
This is why I keep arguing that coherence is not a governance problem in the usual sense. It is a visibility problem first. You cannot govern, audit, permission, or protect what you cannot see. Information sovereignty has a prerequisite, and the prerequisite is knowing what you have.
To be precise about scope: sensing is the first layer of a larger stack the book lays out. There are only four places to intervene on incoherence. You can sense it, constrain it, contain it, or price it, making the team that creates a coordination burden bear the cost it imposes on everyone else. Above all four sits human judgment, reserved for what the lower layers surface. Nadella’s boundary will eventually need the whole stack. But sensing comes first by necessity, not preference: you cannot constrain, contain, or price what you cannot detect.
Build that layer, or the boundary is decoration.
Where to spend the money
The last gap is a budget question, and it is the one that will actually decide whether any of this gets done.
Trust boundaries are not free. Private evals, tenant-bound training environments, a genuinely model-agnostic orchestration layer: these are real investments in engineering and in organizational discipline. No enterprise can build them around everything. Any advice that implies otherwise will be ignored by the people who have to fund it, and they will be right to ignore it.
Narayanan and Kapur draw a line that helps here. They separate the value AI creates from the value anyone manages to capture. The value created will be vast. The open question is who keeps it. Apply that line one level down, inside your own firm. Your people create judgment-value every time they correct the machine. The only question that matters is whether you capture it or the vendor does.
Spend where the exhaust encodes judgment you could not replace and where the deployment is deep enough that its traces map how you actually work. Tolerate leakage where the task is a commodity, the failure is cheap.
The two axes tell you where to spend. Not all incoherence is worth preventing, and by the same logic, not all leakage is worth stopping. Spend where the exhaust encodes judgment you could not replace and where the deployment is deep enough that its traces map how you actually work. Tolerate leakage where the task is a commodity, the failure is cheap, and the trace tells a competitor nothing they do not already know.
An enterprise that hardens every boundary equally has misread the problem exactly as badly as one that hardens none. The first will spend itself into paralysis. The second will donate its judgment one correction at a time, and never see the invoice.
What it does not mean
One caution, because this argument is easy to overcook and the overcooked version is wrong.
Keeping important work away from AI is just retreat dressed as strategy, and it loses. Abstention just gets you slower, with none of the benefits.
The lesson is not “keep your important work away from AI.” That is a retreat dressed as a strategy, and it loses. A competitor who brings AI to their hardest, highest-judgment work, in the posture that work allows, assistance where verification is weak, autonomy where it is strong, and does it inside a proper boundary, gets two things you do not: the compounding and the protection. Abstention gets you neither. It just gets you slower.
There is a real cost to engagement, and Narayanan and Kapur name it. Leaning on a vendor’s AI can erode your unaided skill while building a vendor-specific dependence, a lock-in that works through your own people rather than your contracts. But that is a cost of careless engagement, not of the engagement itself. The answer is the same one the whole piece has been building toward: engage hard, own the loop, keep the orchestration model-agnostic so the skill you build is yours and portable. Abstention avoids the behavioral moat only by forfeiting the capability, which is the worst trade on the board.
There is a reason to move now rather than later. The moat Narayanan and Kapur describe is not yet built. Enterprises have so far been reluctant to feed their material into the flywheel, and the orchestration layer is still, for the moment, thin and swappable. That window does not stay open. The time to build the boundary is before the lock-in compounds, which is to say now.
In consuming intelligence, you are creating intelligence, and what you create should belong to you. The goal is not to stop feeding the machine. It is to make sure the loop closes inside your own walls.
Nadella has the emphasis right. In consuming intelligence, you are creating intelligence, and what you create should belong to you. The goal is not to stop feeding the machine. It is to make sure the loop closes inside your own walls, so that the judgment you spend every day encoding accrues to you instead of leaking to the firm that sold you the model.
That is the difference between an enterprise that compounds and one that is quietly farmed.
That question, whether the judgment you encode every day accrues to you or leaks to whoever sold you the model, is the subject of my book, Coherence, and of everything I am writing here between now and launch. If you are deciding where the boundary has to be hard and where to let the exhaust go, the one-page tool behind the two axes in this piece is the first thing I send when you join the list at coherise.com.