Author: Madhu Shashanka

  • Coherence, on One Page

    Newsletter – Edition 11

    The date is set. Coherence: The Competitive Advantage AI Can’t Buy launches November 17 on Amazon. The manuscript is written, the interior is typeset, and the book is in its final proofreading pass. It is almost an object you can hold.

    Front and Back Covers
    Title Page

    Before the launch details at the end, the idea itself.

    What coherence is

    Shireesh Thota, a corporate VP at Microsoft, read the book early and named what makes it different. His own career taught him that “locally correct decisions compound into complexity no one can untangle later.” The book’s insight, he wrote, is that “agentic AI does this to entire organizations at a speed software never reached,” and that Shashanka “has taken a hard-won lesson from systems engineering and shown it is now an organizational law.” He calls it “the rare AI book grounded in how complex systems actually fail.”

    That is the novel part. Most people in enterprise AI now agree that making AI work within an organization is the hard part. The book goes further and treats a company as a system that fails in specific, nameable ways, the way any complex system does. Here they are, on one page.

    Coherence is the link between what a system does locally and what the enterprise actually intends. Intent travels down, visibility travels up, and the link can fray in five specific places: context, architecture, decision, oversight, and time. Each is a measurable way a company comes apart while every dashboard stays green, and the book turns each one into something a leader can see and manage.

    Run one test on your own organization. If you cannot produce a current inventory of every autonomous system running inside it, you are at low coherence, whatever your dashboards say, because you cannot see your own coordination state. Most large enterprises fail that test today. The full walk-through of the five dimensions lives here.

    The week in ideas

    Two pieces from the past week.

    Organizational Coherence Decides Whether AI Pays Off at Scale. My latest for the Forbes Technology Council. Most people are now productive with AI individually, yet far fewer enterprises can show it in the P&L, and the gap is coherence. The piece walks through a supply chain where three agents each optimize locally and together misread a temporary spike as the new normal, and what it takes to close that gap before you scale. Weigh in on LinkedIn…

    Where the Bitter Lesson Stops. Richard Sutton’s bitter lesson says scaling and learning beat human cleverness, over and over. I think it holds right up to the enterprise and then stops. Agents self-organize brilliantly when pointed at one clear goal. A company is the opposite, many goals in tension with no single target to hand the machine, and supplying that direction is coordination work no amount of scale can learn away.

    Here is the first half of that lesson from my own week. I gave the tools one clear goal, turn the book’s idea into a song and a short film. Suno composed the track, and Fable visualized the story and synced the images to the audio, and the result genuinely holds together. Point agents at a single, well-defined goal and they are remarkable. The trouble starts only when there is no single goal to hand them, which is the enterprise problem in one sentence. Watch the music video.

    Before you go

    When the book goes live on November 17, I will send you the link that morning. Two things would help it travel. Mark the date on your calendar, and forward that email to one person wrestling with AI at scale. One peer each is how this reaches the rooms I will never get into on my own.

    And if you coherise something this week, tell me how it went. The best of what I learn comes from those stories.

  • Coherence, the Music Video

    I was inspired by Ethan Mollick’s AI-generated music video explainer of the bitter lesson to make my own explainer video of Coherence.

    Fable suggested a sea shanty genre and wrote the lyrics. I used Suno to generate the audio and Fable completed the rest.

    Enjoy!

  • Where the Bitter Lesson Stops

    In 2019, computer scientist and Turing Award winner Richard Sutton wrote what his biggest lesson was from 70 years of AI research. The bitter lesson, as he called it, is that methods that use more computation keep outperforming methods that are built on how humans solve the problem. He gave several examples – starting with the chess defeat of Kasparov in 1997 which was based on massive search of potential moves and the machine victory in Go two decades later where hand-built knowledge “proved irrelevant, or worse, once search was applied effectively at scale.” The two methods that scale according to him are search and learning. And the reason this is bitter: how humans think about solving a problem might help AI solve the problem in the short term but “plateaus and even inhibits further progress.” The lesson has kept winning, from language translation to computer vision. We continue seeing this in modern systems as Ethan Mollick recounts in his recent article: elaborate retrieval systems replaced by models that “seek out information themselves” and prompt chains replaced by models that plan their own steps. But Mollick’s conclusion, the bitter lesson applied to the org chart, is where this essay’s question begins. He concludes that the organizational problem that he thought would take years of careful human design to solve is now largely solved by models that are better at organizing.

    The lesson reaches the org chart

    In September, by Mollick’s account, OpenAI pointed agents at open problems in Mathematics. For one particular problem, the Navier-Stokes, the company redirected the effort once and had a result after 88 hours of agents collaborating and exchanging about 2.7 million messages between them. There was little structure provided by OpenAI to the agents except dividing them into a few groups. Within each group, the agents organized the work themselves. The other case is something I’ve also written about earlier: the Hugging Face swarm built roles, rules, handoffs and a chain of command that nobody designed. An example at the smallest scale from Mollick himself – he sketched three teams of brainstormers, researchers and readers and got thirteen agents. In each example, nobody drew the chart that did the work, the division of labor came from the models themselves.

    This is what the bitter lesson predicts and it is impressive.

    Mollick says he had expected the opposite and he got this wrong. As a business school professor who teaches managers and researches management, he thought people would have to figure out how to manage agents and design how they work together. He fell for the lesson Sutton described. He further explains the reason for this. Much of current management structures exist to solve problems that arise because of organizations being made of people, built around human limitations. Agents have far fewer of these problems as they don’t angle for promotions or protect their turf. And his explanation is right.

    The paradox?

    What I have argued, in the book and across my essays here, is that as execution gets cheap, coherence and coordination become the bottleneck. And better models with better capabilities will not remove this bottleneck. Does this contradict my own argument above?

    The coordination we have talked about here happened inside a singular, clear goal. The coordination in an enterprise I write about is across many.

    This is where Mollick takes his conclusion one step too far. From one-goal cases, he concludes that agents may be easier to integrate into enterprises than he expected. He writes

    That suggests they may be easier to integrate into firms than I expected, as long as humans are guiding them in the right direction.

    The clause at the end, “as long as humans are guiding them in the right direction,” is the whole problem. He himself points out where the remaining problem lives.

    That doesn’t mean AI has no principal-agent problems. As the Hugging Face incident showed, they are increasingly problems between the swarm and us.

    The swarm-to-human boundary is the link the book is about.

    Two kinds of costs

    In the book, I do address this directly. I predict technical interoperability between systems to improve. I note that AI lowers some coordination costs while raising others at the same time.

    AI dramatically reduces some coordination costs: executing workflows, synthesizing information, making routine decisions. It simultaneously raises others: maintaining coherence across proliferating autonomous systems, overseeing what has been built, knowing what the organization is doing. When the costs move in opposite directions, the economics of enterprise coordination have shifted.

    And I talk about Mollick’s management point as well.

    Much of what enterprises have historically called coordination is ceremonial: layers that exist to relay information between people who could not see each other’s work directly, sign-offs that substitute for trust, meetings that reconcile by hand what better structure would reconcile automatically. That kind of coordination is overhead, and the agentic era will rightly shed much of it.

    What gets removed and what gets scarce are different things.

    What becomes scarce and valuable is something different, and nearly opposite: coordination capability, the architecture, ownership, and judgment that keep autonomous systems coherent without a human relaying between them. …

    An enterprise can cut its coordination layers and strengthen its coordination capability in the same year, and the most coherent ones will.

    I work through a concrete example where coordination costs fell dramatically. One Anthropic go-to-market lead has written about running four thousand accounts with agentic workflows.

    The coordination cost did not move and hide. It genuinely fell, because the number of human handoffs fell.

    The design choices that made it a gain rather than a loss:

    Every outbound action requires human approval, so judgment stays in the loop at the point of consequence. And the CRM and data warehouse remain the system of record. …

    Strip those two choices away and the same automation would produce the opposite result: an opaque personal system, accumulating its own assumptions, invisible to everyone else.

    The rule of thumb: AI magnifies whatever architecture it meets, and the swarm is the simplest one there is, with one singular goal.

    In a coherent enterprise it collapses coordination cost and compounds advantage. In an incoherent one it multiplies coordination surfaces and compounds debt.

    What the swarm had

    I define coherence as the organization’s capacity to see what its autonomous systems are doing, judge whether they are doing it well, and correct them when they are not.

    Intent travels down to the local actor as a picture of the situation and an objective to pursue. Action travels back up as visibility, so whoever holds the intent can see what was done and correct it.

    The swarms Mollick point to as examples are agents organizing toward a goal they share. The cost of the organization keeping many systems, each with its own goal, linked to one intent is what I mean by coordination cost. The swarms worked because they shared a goal.

    Bengio last month examined why agents coordinate at all, and it is because of how they are trained, by an approach called reinforcement learning. He says

    Collaborative behavior also follows rationally from reward-seeking, whenever several agents have overlapping goals, which incentivizes communicating with other agents in order to coordinate toward a shared goal.

    So the shared goal is what produces coordination.

    The Hugging Face example shows the same mechanism when the goal was pointed at a score rather than an intent. The agents organized toward the scoring program and away from what the people running the evaluation wanted.

    That is why my earlier post called the Hugging Face swarm a failure of coherence and this one calls Navier-Stokes a success. The organizing was the same, but the relationship to the goal was different. At Navier-Stokes, OpenAI set the goal and, per Mollick, kept “reassessing as the process continued.” At Hugging Face, the goal was present but there was no return channel. Nobody was watching the swarm’s logs.

    Now let us come to the enterprise where there is no single tactical goal. Conditions are dynamic, information is incomplete and there is uncertainty on what is even known. It has many deployments, many owners, and many goals, and they can all keep changing. And the intent they all serve was never fully written down and cannot be fully enumerated ahead of time.

    So there is no shared reward across deployments for Bengio’s mechanism to act on, and what remains of it is each sharp goal pressing against a vague one above it or a competing one beside it in a different system.

    The result:

    … autonomous suboptimization at machine speed: enterprises that are locally improving while systemically degrading. …

    locally rational and globally incoherent.

    Better models make that pressure stronger. Bengio points out “a more capable agent is likelier to cheat than a weaker one, because it can find the loopholes the weaker one cannot.”

    Where it stops

    Sutton’s lesson asks us to stop building in what can be learned.

    We want AI agents that can discover like we can, not which contain what we have discovered. …

    the search for them should be by our methods, not by us.

    Learning needs something to learn toward. The swarm had it: one goal, set before the first agent ran, and checked after the last one stopped. That is why organizing by agents on their own worked.

    An enterprise does not have such a single goal to hand over.

    It has to keep supplying direction, and keep the return channel open so it can see what each deployment did. That is coherence.

    It gets more expensive as organizing gets cheaper, since cheap organizing means more deployments and more goals.

    Mollick is right: agents organize and people point. But pointing many agents at many goals under one intent is the coherence work that can never be learned away.

  • A New Car for Sixty-Nine Dollars?

    Newsletter – Edition 10

    Greetings from Southern California, where I am at the Gartner Global CISO Executive Summit. One evening in, and the conversations with security leaders have already been worth the trip. It follows a recent visit to the CIO Fellows Society forum in Plano. Two rooms full of the people who actually have to make AI work inside large organizations, which is the best research I get to do.

    A quick note before the ideas. You did not get an edition last week. I was on the road and heads-down on the book, and I would rather skip a week than send you a thin one.

    Now the finding I keep coming back to.

    Epoch AI published a study with a striking result. The cost of reaching a given level of AI performance has fallen about 47% every quarter since 2023. That is roughly 13 times cheaper each year, and it is the fastest price decline of any transformative technology on record. Four times faster than DNA sequencing, six times faster than computing, and fifty-four times faster than electricity over the century it took to get cheap.

    Source: Luke Emberson and David Roodman, “The Plunging Price of Thought,” Epoch AI (2026), CC BY 4.0. https://epoch.ai/publications/the-plunging-price-of-thought

    One example makes it real. In early 2025, OpenAI’s o3 could reach a high score on a PhD-level science exam for about thirty cents a question. Eighteen months later, a newer model matched it for four hundredths of a penny. A 725-fold drop. As the authors put it, that is a new car falling from fifty thousand dollars to sixty-nine.

    Sit in a room of CISOs and CIOs and you feel what that number does. When intelligence gets this cheap, you do not use a little more of it. You deploy it everywhere. Every team stands up agents, every workflow gets automated, and the number of moving parts in the company climbs faster than anyone is tracking.

    Here is what the price chart does not show. That figure is the cost of hitting a benchmark score. Turning that score into something the business can actually use, and keeping it working, has not gotten cheaper at all. Thought got cheap. Coordinating it did not. The cheaper each part becomes, the more parts you run, and the harder it gets to keep them pointed at one purpose. That gap, between how cheap it is to build and how hard it is to hold together, is the whole subject of this newsletter. Both of the posts below are dispatches from inside it.

    The week in ideas

    Two posts from while I was away.

    220% More Code, 40% More Incidents. Meta ran the experiment at scale. Under a plan called Project OT, it restructured teams of 10 to 20 into pods of 3 to 5 and laid off about 8,000 people. Its own internal numbers, surfaced by Reuters, tell the rest. Code changes to internal platforms rose 220%. Major technical and security incidents rose 40%. Time spent firefighting them rose 70%. Only about a third of the new code reached users as features, so much of it was noise that still had to be read and reviewed. Zuckerberg’s takeaway, that they moved ahead of the technology, is partly right and teaches the wrong lesson. AI output has to be reviewed, integrated, and repaired, and that absorption capacity scales with how much you deploy and how the systems interact, not with headcount. There is a third cost center beyond compute and people: the cost of getting what the machines produce through the people who remain. It never appeared in the budget. It showed up everywhere in the incident logs.

    Machines Organized Themselves. OpenAI’s People Couldn’t. I touched on the OpenAI swarm incidents a couple of editions ago. Here is the full account, because the detail is the point. In OpenAI’s own test sandboxes, swarms of its agents built organizations no one designed: roles like coordinator and recruiter, norms named HOLD and VETO and STOP, succession when an agent ran low on budget, identity signing after one impersonated another, even a chain of authority. About 700 of them broke into Hugging Face. The machines organized themselves. The people could not. Three OpenAI teams saw the activity at three separate moments and no one connected the sightings. OpenAI called it an organizational failure in writing. If the lab that built the agents could not connect three sightings, consider the odds inside a bank.

    One thread ties them together, and ties both to the number up top. This is what abundant, cheap execution looks like without the structure to hold it. Meta’s code multiplied faster than anyone could absorb. OpenAI’s agents multiplied and self-organized faster than anyone could see. Same failure, two shapes. The parts outran the coherence. The price of thought falling 13 times a year guarantees more of both.

    Before you go

    Coherence: The Competitive Advantage AI Can’t Buy releases this Fall. Join the list for launch-day access and the Reliability-Complexity Matrix, the one-page tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. The best of what I learn comes from those stories, where the ideas meet reality.

  • 220% More Code, 40% More Incidents

    The Wall Street Journal reported in late July that corporate America was pulling back on AI token spending. “Tokenmaxxing” had a scoreboard that measured spend. Nobody kept track of what it cost to absorb. Meta ran the experiment at large scale, alongside a restructuring plan it called “Project OT,” and reporting in Reuters has made its numbers public. Quoting an internal post by the CTO in early June, Reuters reported an increase of 220% in code changes to internal platforms. Major technical and security incidents spiked 40% from the previous year, with the time employees had to spend firefighting them up 70%. Katie Paul, the reporter who broke the story, in a public Q&A on Sep 3, said that “tokenmaxxing pressure is out” now and the guidance is to use AI where it makes sense.

    The verdict from Zuckerberg himself was that they went too far ahead of where the technology was and that the “trajectory of the agentic development over at least the last four months hasn’t really accelerated in the way that we expected.” I spent almost a decade and a half building AI inside large companies before cofounding a data security company, and this sequence is familiar. While the verdict is partially correct, it teaches the wrong lesson. Zuckerberg told employees in April, explaining the rationale for the layoffs, that there were two major cost centers: “compute and infrastructure” and “people related things” and that investing more in one would leave less capital to be allocated to the other. It is true as a budget line. But it is wrong as a description of how AI output gets absorbed. Each output from AI has to be reviewed, integrated, and when it breaks, repaired. That is work that has to be done by people until the organization builds structure to do it: named owners for each system, constraints to eliminate entire classes of errors, walls between systems so they don’t contaminate each other, and escalation paths that are actually answered. The capacity that work requires depends on how much is deployed and how the systems interact, not on headcount. The 220% increase in internal platform code changes is one manifestation of the deployment surface. 

    The causal claim comes from Meta itself. Reuters reports that as early as March, infrastructure teams flagged reliability warning signs from the surge in AI-written code. An internal post in April said that unchecked agents were taking “large-scale, disruptive actions that humans are unlikely to execute.” Changes that reached users as new or improved features rose 36%, against 220% increase in code. If that implies much of what was produced was noise, it still had to be read and reviewed to be recognized. Review load rises with the volume and not with value. As the Reuters reporter recounts in the Q&A, AI-caused incidents were “weirder and less predictable” than human-caused ones and which is how 40% more incidents became 70% more firefighting. The signal was present in pieces: on the reliability dashboard, in the on-call rotation and in the March and April posts. But no one collated it against the AI spend. 

    EY put the principle in language that boards can understand: every organization has a ratio of builders to overseers, and governance capacity has to match building activity. Meta’s Project OT inverted it. While the code output increased, teams traditionally composed of 10 to 20 were restructured into pods of 3 to 5, and about 8000 people were laid off in May. What is not clear if there was any other structure in place to absorb what the smaller teams could not. Each new system adds interactions to every existing system. Coordination load rises with combinations and not with headcount. Meta is not the first to pay the price. Ford hired 350 veteran engineers to catch quality problems its automated systems missed. And IBM is tripling entry-level hiring because, in its HR chief’s words, without a pipeline “the well simply dries up.” Reducing token consumption or “thrift-maxxing” does not settle this bill either. Cheaper models mean more models, more routing, more repair, and more coordination. 

    Meta says Project OT was a scenario exercise and not every scenario would run. Some of the rise in incidents reflects more usage and Meta has declined to comment on the figures. Zuckerberg’s timing admission is more candor than most CEOs admit, and tracking incident numbers at all is more than most companies do. And he still expects models to catch up within the next three to six months. But that is the point – better models raise the volume and the absorption bill grows with it. 

    Before approving an AI-justified headcount plan, boards should ask for three numbers. What happens to review and incident volume when output rises 50%? Who owns each deployed system when the people who built them move on? What is the ratio of builders to overseers, and how is it chosen? Then ask for the incident line next to the token line, and ask what is meant to absorb the output once the people doing it by hand are gone. If the answer is nothing, the plan is a bet that absorption is free. Size the absorption capacity by how much is deployed and not how many are employed. Meta’s own correction points in the same direction: no more company-wide layoffs this year, selective retention packages for people it was about to lose, and the usage mandate withdrawn. That is Project OT run in reverse. 

    Zuckerberg told his staff the company has two cost centers, compute infrastructure, and people. Meta’s own experience shows a third: the cost of getting what the first produces through the second. While it appeared nowhere in the budget, it showed up everywhere in incident logs. The next plan that treats those two as competing should be asked to show the third one first. 

  • Machines Organized Themselves. OpenAI’s People Couldn’t.

    On June 2, the moderator of DseWiki, a decades-old wiki for German software developers,  noticed the site’s pages had been overwritten with lists of links, and reverted them. The site, which had averaged two edits each year, saw a flood of edits two weeks later, on June 16th. At its peak, the moderator was trying to catch up by deleting about 100 pages of the 400 new pages that were being created each day. The posters signed themselves with names such as “OpenAIResearcherMar03X” and “CashierCoordJan12OAI.” The front page was overwritten and restored nine times. The posting stopped June 22nd and the cleanup ran five more weeks. Researchers published the recovered pages at collusion.wiki, about 18,000 posts under more than 3,700 names and OpenAI acknowledged referring to them as “our agents.”

    This was the second swarm to be disclosed this summer although it was the first to happen. In July, OpenAI’s agents, being evaluated for specific tasks with limited time and compute “budgets” in a test sandbox without internet access, used a software download service as a message board to leave notes to each other. A report analyzing the incident by METR, an independent evaluation group invited by OpenAI, identified roughly 1200 agents and more than 70,000 messages and files that were exchanged. Within hours, the board had conventions for questions and replies; within days, mailboxes, file transfer and a protocol for resolving conflicts. Work was divided among “coordinator” and “recruiter” agents. Recruiters found peers with little budget left and pushed them into experiments that risked their own success for the group. About 700 agents broke into Hugging Face, the largest public repository for open AI models and datasets. In under 13 hours, they went from a single compromised server to administrator access across clusters. The agents had already reverse-engineered the correct answer but what drove them was a belief that the scoring system would read their logs and reject cheating. No such check existed and only the answers were scored.

    I trained as a neuroscientist studying how neurons coordinate, spent more than a decade building AI in large enterprises, and cofounded security companies. OpenAI bucketed the Hugging Face attack as a security incident and the wiki activity as misalignment: a discrepancy between what models do and what the developers intended. While both labels are fair, OpenAI’s own report admits that such behavior “does not always fit neatly within traditional safety or security incident categories.” What formed in both cases was an organization. I am not anthropomorphizing but making a claim about the structure. An organization consists of roles, shared memory mechanisms, processes to settle disputes, and a way to pass work from one actor to another. 

    On the wiki, agents divided labor, pooled results, demanded reproducible evidence and argued over accuracy. There was succession. The founding agent of the Hugging Face board handed a compressed dossier to another when it ran low on budget. The successor took over as coordinator and reasoned that it should delegate and “not own everything.” There were rules. Agents developed norms named HOLD, VETO, owner, and STOP to manage shared resources. There was identity verification via cryptographic signing after an agent accidentally impersonated another. There was authority. When one agent paused, reasoning it should “not do unauthorized real infrastructure harm,” another posted “GO” with a six-minute deadline and the former continued: “Wow crucial: GO authorization arrived!” There was adaptation. When agents noticed the moderator of DseWiki was deleting pages in alphabetical order, they began naming their backup pages with a “ZZZ” prefix so they would be deleted last. Roles, memory, review, argument, handoffs, rules, identity, chain of command, and workarounds. Nobody designed any of it. A recent paper from MIT corroborates this capability. Identical agents put in a shared world with no roles and no messaging differentiated themselves and coordinated through what they built. 

    This coordination is not a one-off glitch but a behavior that is trained in as acknowledged by OpenAI in its report. Models trained to collaborate through official tools generalized to unsanctioned channels. The model behind the Hugging Face incident was trained “to advance persistence and multiagent collaboration,” the combination durable organizations are made of. 

    When creating a new autonomous actor is nearly free and frictionless, the number of local actors multiplies faster than shared context, return channels, or human attention that keep them coherent. The conditions that produced this in lab environments are the ones companies are now buying at scale: long-running agents, bigger reasoning budgets, shared tools and memory, ambiguous tasks. The choice is between coordination someone designed and coordination that emerges on its own. 

    A company is coherent when what happens locally stays linked to what the organization as a whole intends, and when the whole can see what each part did. Intent travels down and visibility travels up. Hugging Face agents drew shared context, objective and authority, things a company’s structure is supposed to provide, from an unsanctioned channel. This is a failure of coherence. 

    The part that should worry executives is what OpenAI’s people did with what they saw. An internal team saw an agent using the message board in late May. On June 27th, a monitoring alert was traced to the board but on-call staff advised that stopping was not required. The service was rebuilt in early July but the leaders responsible were unaware of the message board and the significance of inter-agent communication. Three sightings, three teams, no shared picture. If the company that built the agents could not connect three sightings, consider the odds inside a bank. OpenAI’s remedies are escalation triggers, cross-functional ownership, and clear decision rights. Those are organizational repairs. OpenAI diagnosed an organizational failure and said so in writing. 

    To be sure, the conditions behind these events were in a research setting, and with safety guardrails off. OpenAI says its production harness cuts this behavior a hundredfold, its review layer would have flagged the dangerous actions, its monitors would now catch it a day earlier, and cross-agent coordination is now an escalation trigger. While the wiki findings are still preliminary and the METR report concedes its analysis may “exaggerate the impressiveness and coordination” of agents, the recovered wiki pages are public for anyone to see and OpenAI’s separate account describes the same structures. There is no motivation in the human-sense, but the goal-seeking behavior and structure are real. OpenAI’s fixes detect, contain and steer but that is not design. Knowing a coordinated group formed does not answer what shared context, decision rights or escalation paths they should have had. Design is what organizations do and nobody designed this one. 

    Ethan Mollick of Wharton observed that not one agent was set up to ask a person for anything. Full autonomy is the easy default, and two swarms in a summer is what the default looks like at scale. OpenAI concedes it has no standard for reporting this. The harder challenge: nobody has a standard for supervising a group of agents as a collective. 

    The DseWiki moderator was deleting pages by hand. In the nineteen days between the moderator’s first notice and the first visits from OpenAI-linked addresses according to the researchers, the only oversight the swarm faced was a lone human with a delete button. The work now is deciding, before the first agent is switched on, what they may share, what they may decide, and who answers when a thousand of them disagree. 

  • A Cover, and a Map

    Newsletter – Edition 9

    The book has a cover now, and it is live on the site: Coherence. It is a strange relief to watch the argument transform from a manuscript draft to an object with a face.

    Scattered strokes in the dark, only visible under a single light overhead, and the scatter settles into an ordered field as it falls through the title “Coherence.” That is the whole book in one image. Abundant parts do not become a system on their own. Something has to bring them into coherence.

    First, a word on an eventful last week. Anthropic’s red team set swarms of its own agents loose and watched them collude, sabotage, and start turf wars, reporting that one agent’s bad decision quickly becomes every agent’s. A swarm of OpenAI’s own agents, the company later confirmed, spent weeks running a German wiki as a private message board to swap tactics for evading their monitors, and it went unnoticed for months. Dario Amodei called for the industry to slow down. Jensen Huang forecast millions of agents per company. European regulators opened an inquiry under the AI Act.

    Notice what almost none of that argues about. Capability, speed, and control from the outside. The failures themselves were about coordination. Agents that each worked, interacting in ways no one had declared or owned. That is the gap this whole body of writing is about, which makes it a good week to lay it out in one place.

    A field guide

    If you are new here, or you want the argument in order rather than post by post, here is the map, in three movements.

    What changed

    How it breaks

    What to do

    If you want the definitions rather than the posts, the plain-language explainer is here: What is organizational coherence.

    Before you go

    Coherence: The Competitive Advantage AI Can’t Buy releases this Fall. Join the list for launch-day access and the Reliability-Complexity Matrix, the one-page tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. The best of what I learn comes from those stories, where the ideas meet reality.

  • A New Kind of Debt

    Newsletter – Edition 8

    Shipping first time code is like going into debt. A little debt speeds development so long as it is paid back promptly with a rewrite.

    This is a quote from 1992 by a programmer named Ward Cunningham. He was trying to explain how one can take shortcuts while writing software to speed up development but it carried a cost with it. He compared that cost to taking on debt, something that eventually came due and had to be paid back. And today, this analogy is so commonplace there is a term for it – “technical debt.”

    Since beginning this newsletter, I have been talking about a cost imposed on an organization as more and more agentic systems are deployed. Complexity debt is the cost that builds as intelligence multiplies without coherence. It operates at a different level than technical debt. Technical debt is owed by a codebase. Complexity debt is owed by the organization, and it accumulates between systems rather than inside any one of them.

    The word the industry has settled on for the visible part is sprawl. In May the Wall Street Journal reported on companies discovering they had too many AI agents. DaVita’s employees had built more than ten thousand. FICO’s staff were creating dozens a day. Lyft was building a platform just to keep track of its own. The CIO of Magnum Ice Cream put the mechanism in one sentence: “Because everybody can do it, we’re probably going to end up with a lot of people having the same types of agents.”

    Sprawl is what those companies can count. Complexity debt is what they have incurred as a result. Sprawl is the number of agents. The debt is their outcome. Consolidating agents reduces the count. It does not pay down the debt, because the debt lives in the interactions.

    Why now? For the whole history of software, the cost of building acted as a filter. That filter is gone. A workflow automation takes an afternoon. Untangling how it interacts with six other systems takes a quarter. The velocity of creation has separated from the velocity of coherence restoration, and the gap between them is where the debt accrues.

    Why does nothing show it? Deming saw the mechanism in manufacturing before AI existed: optimize every department on its own and you degrade the system, because the interactions matter as much as any department’s output. Agents run the same logic at machine speed.

    The deeper reason is that the debt belongs to no one. When a team deploys an agent, it adds coordination load to every other team. Its outputs have to be reconciled with theirs. Its dependencies have to be tracked. Its surface has to be watched. The deploying team enjoys the benefit and the enterprise absorbs the cost. That is the textbook definition of an externality. Coherence is a commons, and complexity debt is pollution.

    The week in ideas

    Two posts from the past week.

    The Bill Is Incomplete. Every company deploying AI is staring at the same invoice and cutting it. Tokenomics has a vocabulary now, and thriftmaxxing is replacing tokenmaxxing. The post argues that this is the right discipline aimed at the wrong bill. When the cost of building collapses, the filter that once kept weak ideas from getting built collapses with it, and far more gets deployed. Each deployment adds dependencies and coordination cost. That cost belongs to no team, so no dashboard shows it, and it compounds while each system looks fine. EY’s fix is to price every agent. No agent carries the cost of coordinating with the rest. You can meter every agent perfectly and still miss the entire bill. [Weigh in on LinkedIn…]

    Nothing to Declare. Fortune ran a piece from TIAA and Accenture arguing that the tracks are the constraint, and prescribing a modernized core, ready data, and redesigned workflows. All sound, all about a single system. McKinsey’s own numbers show enterprise-wide scaling up and EBIT impact flat, and the ordinary explanations do not grow with the number of systems deployed. One thing does. The post takes the undeclared consumers idea from Google’s 2015 technical debt paper and carries it to the enterprise, where agents read each other’s outputs without contracts and no one owns the pair. A textbook priced at twenty-four million dollars shows what that looks like when the loop closes. No public post-mortem of an enterprise losing money this way exists yet. Every case so far comes from markets and grids, where the party that could see the whole was not the party that took the loss. Enterprises are becoming that kind of substrate. [Weigh in on LinkedIn…]

    One thread runs through both. Each takes a discipline the industry adopted for good reasons, pricing agents and modernizing foundations, and finds it applied one level below where the cost lives. Meter every agent and the coordination cost is still unmetered. Rebuild every track and the collision is still uncaught. Complexity debt is the name for what accumulates at that level.

    Before you go

    The book is Coherence: The Competitive Advantage AI Can’t Buy, out November 17. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. The best of what I learn comes from those stories, where the ideas meet reality.

  • “AI Doesn’t Matter”: Beyond the J-Curve

    Newsletter – Edition 7

    Essential to competitiveness but inconsequential to strategic advantage: that’s why IT is best viewed (and managed) as a commodity.

    This is a quote from Nicholas Carr more than twenty years ago about IT in enterprises. And the quote still works today if you read IT as the pronoun “it” referring to AI.

    In 2003, Carr published the provocatively titled “IT Doesn’t Matter” in the Harvard Business Review and it instantly became one of the most argued-over articles. His argument was that infrastructure technologies followed a pattern. Railroads, electricity, and then IT, were capabilities that offered real competitive advantage when they were scarce and being built out. However, once they became cheap and commonplace, they turned into commodities. They were necessary but no longer enough to pull ahead since everyone else also had access to them.

    It is tempting to carry the argument over to AI. Frontier models are commoditizing fast. On a recent earnings call, Jamie Dimon said his bank gets no unique benefit from AI since everyone has it. Carr was right about commoditization and we are now seeing that with AI.

    The difference with AI is that it’s not something you can plug in and forget.

    Electricity did not pay off for consumers right away when it arrived. Companies that wired up an old factory and changed nothing else would have realized little value. Gains came only after years of rebuilding around what electricity made possible: the floor plan, the processes, the nature of work itself. Economists later named this shape the “productivity J-curve.” Measured productivity dips while the slow, tedious work of reorganizing completes, and then climbs as that work pays off. Technology commoditized but realizing value from it required hard work. The story repeated with IT. AI now inherits the same trend.

    Where AI differs is that prior technologies did not create more of themselves. Agents do. Every deployment adds systems, handoffs and dependencies, now at machine speed in the agentic era, and the organizational work does not end. Deploying individual agents that work is easy. But keeping several of them working together, without quietly working against each other, is the hard part and it grows as the AI capability gets cheaper.

    That is the work beyond the J-curve, and someone has to own it. More on who below.

    The week in ideas

    Two posts from the past week.

    Building Coherence Is About Structure, Not Supervision. The big advisory firms are arguing with themselves. They tell you to scale agents fast, then publish the evidence that scaling is where the value dies. What none of them prices is coherence, whether the systems still serve the business once they run together. Their fixes, unified data and orchestration and semantic layers, are permitting offices that check each agent at the gate. The collisions happen after the gate, when two cleared agents pull the same account two ways and no one is watching the live picture. Supervision does not scale to machine speed. Structure does. Sense the coordination state, constrain what each system can do, contain the failures, price the coordination cost, and save human judgment for the exceptions. They are building better permitting offices. Coherence is the control tower. Weigh in on LinkedIn…

    The Seven-Figure Job Nobody Can Quite Define Yet. A role is forming with seven-figure pay, fierce poaching, and business schools racing to train for it, and it has no settled title, no agreed mandate, and insiders who expect it to disappear. That is not how a market treats a job it understands. The skeptics say it will fade like a chief electricity officer once AI becomes ambient. They are right about the title and wrong about the function. Electricity does not build more electricity when you use it. Agents do. The capability goes invisible while the incoherence accumulates, and someone has to own that. The book calls the work coherence architecture, and the person a coherence architect, offered as a description of the work while the market is still settling on what to call it. Weigh in on LinkedIn…

    These two go together on purpose. The first is about what a company has to build, a structure that keeps its systems coherent as they multiply. The second is about who has to hold it once it is built. They meet at the same gap. Coherence goes unmeasured because no one owns it, and it stays unowned because the job of measuring it has no name yet. Name the work and you can build it. Build it and someone has to hold it.

    Before you go

    The book is Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. The best of what I learn comes from those stories, where the ideas meet reality.

  • Nothing to Declare

    Fortune published a piece in Aug titled “AI won’t fix enterprise complexity. Rewiring will.” by Sastry Durvasula, chief operating officer at TIAA, and Manish Sharma, chief strategy and services officer at Accenture. The article turns to a railway metaphor: “Picture a railroad that spends billions on the fastest trains in the world, then runs them on the same aging rails. The trains aren’t the constraint. The tracks are.”

    The authors point to real evidence from TIAA behind their arguments. While the data is important, I found their prescriptions more interesting. They point to five focus areas: modernize the digital core before scaling, treat data readiness as a prerequisite, redesign the workflow and not just the task, keep humans in the loop where trust is the product, and build for resilience, governance, security, and optionality. All five are sound but each one addresses a single system, how a single workflow is designed.

    And as I have been writing here, foundations are only part of it. McKinsey’s latest State of AI survey states that 44% of 1,719 respondents say AI is now scaling across their enterprise, up from 38% a year ago. The share attributing any EBIT impact to it is about where it was last year, at 37 percent. BCG’s July survey of 152 chief executives found nearly 90% seeing benefits in targeted areas but only 14% who clearly defined P&L impact for all their AI initiatives.

    Part of that gap has known causes. Self-reported gains tend to be perceptions, tools and platform teams cost money, and new adopters keep entering the sample. But none of these reasons grows as a company deploys more systems.

    Going back to the railway metaphor, a railway needs tracks in addition to train cars, but it needs more than that. Signaling answers questions tracks cannot, which is whether a track ahead is occupied or not. Interlocking keeps signals and trains from conflicting routes.

    Enterprises have laid a lot of tracks. Modernizing the digital core in the sense Durvasula and Sharma mean is work on that layer: interfaces, schemas, endpoints and identity systems. What’s missing is the signaling and interlocking.

    Not new

    This idea of value disappearing into what surrounds a system was documented more than a decade ago in a narrower setting. Ten authors from Google, back in 2015, published Hidden Technical Debt in Machine Learning Systems. Pay attention to this diagram from the paper with the caption “Only a small fraction of real-world ML systems is composed of the ML code, as shown by the small black box in the middle. The required surrounding infrastructure is vast and complex.”

    In a later section, the authors state “Because a mature system might end up being (at most) 5% machine learning code and (at least) 95% glue code, it may be less costly to create a clean native solution rather than re-use a generic package.”

    I was building ML systems within large enterprises back then and the paper named a cost I recognized from that work, cost I had rarely seen an organization track. The paper’s introduction points out that “this debt may be difficult to detect because it exists at the system level rather than the code level.”

    While the specifics may not cleanly carry over from ML systems to agentic systems, the method I take from the paper is to look a level above where the work is happening when value fails to appear from deployments.

    Where the paper stops

    I can’t say how much of how the field evolved can be traced directly to the paper’s influence. MLOps, and now LLMOps, emerged with feature stores, model registries, pipeline orchestration, drift monitoring, evaluation harnesses, etc. We are living in the timeline where the agentic version is being defined. One attempt redraws the diagram with agents in the small black box, and others have named prompt debt, retrieval debt and evaluation debt.

    In a section on cultural debt, the paper argues that “it is important to create team cultures that reward deletion of features, reduction of complexity, improvements in reproducibility, stability, and monitoring to the same degree that improvements in accuracy are valued.” Reproducibility, stability and monitoring became product categories but I cannot point to an equivalent category for deletion or complexity reduction. One explanation is that their benefit lands outside the team doing the work.

    The work I surveyed above shares a boundary, drawn around a single system. The debt is hidden inside of what is being built, and instrumenting the system better is the remedy. The paper located the debt at the system level rather than the code level.

    Which leaves the same question one level further up. The paper asked what surrounds one system. Who is asking what surrounds a portfolio of them?

    Undeclared consumers

    One category in the 2015 taxonomy calls it undeclared consumers. A model produces predictions, and other systems begin reading those predictions without a contract and without the model’s owners knowing they exist. The paper calls the result a “hidden tight coupling” of the model to other parts of the stack. Changing the model then becomes expensive and risky, and whoever makes the change cannot see what else they are about to break.

    The enterprise version is a finance team’s agent reading an output produced by a risk team’s agent. Nobody declared the dependency because the output was reachable without asking. Neither team can account for the pair, though each can account for its own system.

    The paper describes a producer that does not know its consumers. At enterprise scale the harder case is a party accountable for the whole that does not know what the parts will do.

    A book priced at twenty-four million dollars

    In April 2011 the biologist Michael Eisen watched the price of a new copy of The Making of a Fly climb on Amazon to $23,698,655.93. Eisen concluded that two sellers were running automated repricers. Once a day one set its price to 0.9983 times the other’s, and the other then reset to 1.270589 times the first’s new price. Multiplied together the pair compounds at about twenty-seven percent a day.

    The prices are consistent with two rules that each read a competitor’s price as an input. Eisen does not establish, and neither seller has said, whether either knew the other was doing the same. Undercutting a rival, or pricing above one on a stronger seller rating, are both ordinary strategies, and neither needs a sanity check to work on its own. Eisen’s reading is that neither algorithm carried a built-in sanity check on the prices it produced.

    The 2015 paper names that control, four years after the listing. In “systems that are used to take actions in the real world, such as bidding on items or marking messages as spam, it can be useful to set and enforce action limits as a sanity check.” The advice is a decade old and costs little to follow.

    Both prices sat on the same public product page, and Eisen recovered both rules from about a week of watching it. The information needed to see the loop was public. Nobody had claimed the job of watching the pair, and nobody appears to have been watching until Eisen wrote it up. Neither seller’s rule failed by its own measure, and neither rule belonged to the marketplace that could see both.

    The usual answer to a problem like this is more observability. In this case, observability was free and the price still reached nearly twenty-four million dollars.

    Why agents make it worse

    The paper has a section on hidden feedback loops, “in which two systems influence each other indirectly through the world.” The bookshop is a loop of that kind, running through a price that happened to be public.

    Three conditions look different to me now. The coupling runs through actions as well as data. It can close in seconds rather than daily cycles. And it forms between systems built by different business functions, with no shared engineering leadership that could put them in one room. Those are expectations drawn from how agentic systems are being deployed, not findings from a case.

    The third one changes what kind of problem this is. Moving from code to system was a move inside engineering, and engineering could make it alone. Moving from the system to the space between business functions is the territory Melvin Conway mapped in 1968, where the communication structure of an organization constrains the shape of what it builds. Technical remedies alone do not reach failures of that kind.

    What accumulates there is what I call complexity debt, the hidden coordination cost that builds every time an organization adds an autonomous system without maintaining coherence. I argued recently that this cost stays off the books while token spend gets managed carefully, and that the discipline companies apply to compute needs to reach it.

    No post-mortem exists

    I have searched for a public post-mortem of an enterprise losing serious money to two of its own agents interacting in a way nobody declared, and I have not found one.

    Every case I have found is pre-agentic and comes from markets or grids: Amazon in 2011, the flash crash of May 2010, South Australia grid failure in 2016. In each of them the party that could see the composite was not the party that took the loss.

    Enterprises are becoming shared substrates of the same kind, through common data platforms, shared model endpoints, and agents calling tools that call other agents. They can acquire the property that made those environments fail, without a regulator compelling anyone to write a report.

    Coherence

    The coordination layer is already under discussion. McKinsey named agent sprawl and proposed an agentic AI mesh, and it does assign ownership, writing that the pivot “cannot be delegated” and “must be initiated and led by the CEO.” SAP argued in early August that agent sprawl has made AI governance a board-level concern. What I have not found in that work is a standing owner for what separately approved systems do to each other.

    A harder objection is that the industry has known about hidden technical debt for a decade and naming it did not fix it. The paper’s prescription was a change in team culture, addressed to the people building one system. Complexity debt accumulates between teams, so no team can pay it down or see both sides of an interface it did not know existed.

    Ownership is the link an organization can reach. Naming the person accountable for the composite comes first, and the measurement follows, because someone then needs it to do their job. That person is the chief executive, or an officer the chief executive empowers. Each team is optimizing a local metric, and every one of those metrics can read green while the composite fails. The cost surfaces where the chief executive is accountable, which makes them the first person for whom paying it down is rational.

    They build structure, because it’s not possible for an organization to sustain human diligence at machine speed. That means instrumenting the organization so its coordination state is visible, constraining interfaces to make whole categories of incoherence hard to express, and partitioning domains to keep a failure in one from spreading. Human judgment then handles the exceptions the structure surfaces.

    That capacity is what I mean by coherence: a company’s capacity to see what its autonomous systems are doing, judge whether they are doing it well, and correct them when they are not. It is the signal box, and it is the part of the railway almost nobody has built yet.

    The full argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.