Essays on how companies hold together as they fill with abundant intelligence. Written for CEOs, enterprise leaders and executives navigating the organizational impact of AI and agentic AI.

AI and agentic systems are changing organizational design, operating models and the way companies work. The problem isn’t simply redesigning the organization for AI. It’s keeping the redesigned organization coherent as AI accelerates complexity.

Looking for something else? My academic publications and my Concentric AI writing live elsewhere.

Nothing to Declare

Fortune published a piece in Aug titled “AI won’t fix enterprise complexity. Rewiring will.” by Sastry Durvasula, chief operating officer at TIAA, and Manish Sharma, chief strategy and services officer at Accenture. The article turns to a railway metaphor: “Picture a railroad that spends billions on the fastest trains in the world, then runs them on the same aging rails. The trains aren’t the constraint. The tracks are.”

The authors point to real evidence from TIAA behind their arguments. While the data is important, I found their prescriptions more interesting. They point to five focus areas: modernize the digital core before scaling, treat data readiness as a prerequisite, redesign the workflow and not just the task, keep humans in the loop where trust is the product, and build for resilience, governance, security, and optionality. All five are sound but each one addresses a single system, how a single workflow is designed.

And as I have been writing here, foundations are only part of it. McKinsey’s latest State of AI survey states that 44% of 1,719 respondents say AI is now scaling across their enterprise, up from 38% a year ago. The share attributing any EBIT impact to it is about where it was last year, at 37 percent. BCG’s July survey of 152 chief executives found nearly 90% seeing benefits in targeted areas but only 14% who clearly defined P&L impact for all their AI initiatives.

Part of that gap has known causes. Self-reported gains tend to be perceptions, tools and platform teams cost money, and new adopters keep entering the sample. But none of these reasons grows as a company deploys more systems.

Going back to the railway metaphor, a railway needs tracks in addition to train cars, but it needs more than that. Signaling answers questions tracks cannot, which is whether a track ahead is occupied or not. Interlocking keeps signals and trains from conflicting routes.

Enterprises have laid a lot of tracks. Modernizing the digital core in the sense Durvasula and Sharma mean is work on that layer: interfaces, schemas, endpoints and identity systems. What’s missing is the signaling and interlocking.

Not new

This idea of value disappearing into what surrounds a system was documented more than a decade ago in a narrower setting. Ten authors from Google, back in 2015, published Hidden Technical Debt in Machine Learning Systems. Pay attention to this diagram from the paper with the caption “Only a small fraction of real-world ML systems is composed of the ML code, as shown by the small black box in the middle. The required surrounding infrastructure is vast and complex.”

In a later section, the authors state “Because a mature system might end up being (at most) 5% machine learning code and (at least) 95% glue code, it may be less costly to create a clean native solution rather than re-use a generic package.”

I was building ML systems within large enterprises back then and the paper named a cost I recognized from that work, cost I had rarely seen an organization track. The paper’s introduction points out that “this debt may be difficult to detect because it exists at the system level rather than the code level.”

While the specifics may not cleanly carry over from ML systems to agentic systems, the method I take from the paper is to look a level above where the work is happening when value fails to appear from deployments.

Where the paper stops

I can’t say how much of how the field evolved can be traced directly to the paper’s influence. MLOps, and now LLMOps, emerged with feature stores, model registries, pipeline orchestration, drift monitoring, evaluation harnesses, etc. We are living in the timeline where the agentic version is being defined. One attempt redraws the diagram with agents in the small black box, and others have named prompt debt, retrieval debt and evaluation debt.

In a section on cultural debt, the paper argues that “it is important to create team cultures that reward deletion of features, reduction of complexity, improvements in reproducibility, stability, and monitoring to the same degree that improvements in accuracy are valued.” Reproducibility, stability and monitoring became product categories but I cannot point to an equivalent category for deletion or complexity reduction. One explanation is that their benefit lands outside the team doing the work.

The work I surveyed above shares a boundary, drawn around a single system. The debt is hidden inside of what is being built, and instrumenting the system better is the remedy. The paper located the debt at the system level rather than the code level.

Which leaves the same question one level further up. The paper asked what surrounds one system. Who is asking what surrounds a portfolio of them?

Undeclared consumers

One category in the 2015 taxonomy calls it undeclared consumers. A model produces predictions, and other systems begin reading those predictions without a contract and without the model’s owners knowing they exist. The paper calls the result a “hidden tight coupling” of the model to other parts of the stack. Changing the model then becomes expensive and risky, and whoever makes the change cannot see what else they are about to break.

The enterprise version is a finance team’s agent reading an output produced by a risk team’s agent. Nobody declared the dependency because the output was reachable without asking. Neither team can account for the pair, though each can account for its own system.

The paper describes a producer that does not know its consumers. At enterprise scale the harder case is a party accountable for the whole that does not know what the parts will do.

A book priced at twenty-four million dollars

In April 2011 the biologist Michael Eisen watched the price of a new copy of The Making of a Fly climb on Amazon to $23,698,655.93. Eisen concluded that two sellers were running automated repricers. Once a day one set its price to 0.9983 times the other’s, and the other then reset to 1.270589 times the first’s new price. Multiplied together the pair compounds at about twenty-seven percent a day.

The prices are consistent with two rules that each read a competitor’s price as an input. Eisen does not establish, and neither seller has said, whether either knew the other was doing the same. Undercutting a rival, or pricing above one on a stronger seller rating, are both ordinary strategies, and neither needs a sanity check to work on its own. Eisen’s reading is that neither algorithm carried a built-in sanity check on the prices it produced.

The 2015 paper names that control, four years after the listing. In “systems that are used to take actions in the real world, such as bidding on items or marking messages as spam, it can be useful to set and enforce action limits as a sanity check.” The advice is a decade old and costs little to follow.

Both prices sat on the same public product page, and Eisen recovered both rules from about a week of watching it. The information needed to see the loop was public. Nobody had claimed the job of watching the pair, and nobody appears to have been watching until Eisen wrote it up. Neither seller’s rule failed by its own measure, and neither rule belonged to the marketplace that could see both.

The usual answer to a problem like this is more observability. In this case, observability was free and the price still reached nearly twenty-four million dollars.

Why agents make it worse

The paper has a section on hidden feedback loops, “in which two systems influence each other indirectly through the world.” The bookshop is a loop of that kind, running through a price that happened to be public.

Three conditions look different to me now. The coupling runs through actions as well as data. It can close in seconds rather than daily cycles. And it forms between systems built by different business functions, with no shared engineering leadership that could put them in one room. Those are expectations drawn from how agentic systems are being deployed, not findings from a case.

The third one changes what kind of problem this is. Moving from code to system was a move inside engineering, and engineering could make it alone. Moving from the system to the space between business functions is the territory Melvin Conway mapped in 1968, where the communication structure of an organization constrains the shape of what it builds. Technical remedies alone do not reach failures of that kind.

What accumulates there is what I call complexity debt, the hidden coordination cost that builds every time an organization adds an autonomous system without maintaining coherence. I argued recently that this cost stays off the books while token spend gets managed carefully, and that the discipline companies apply to compute needs to reach it.

No post-mortem exists

I have searched for a public post-mortem of an enterprise losing serious money to two of its own agents interacting in a way nobody declared, and I have not found one.

Every case I have found is pre-agentic and comes from markets or grids: Amazon in 2011, the flash crash of May 2010, South Australia grid failure in 2016. In each of them the party that could see the composite was not the party that took the loss.

Enterprises are becoming shared substrates of the same kind, through common data platforms, shared model endpoints, and agents calling tools that call other agents. They can acquire the property that made those environments fail, without a regulator compelling anyone to write a report.

Coherence

The coordination layer is already under discussion. McKinsey named agent sprawl and proposed an agentic AI mesh, and it does assign ownership, writing that the pivot “cannot be delegated” and “must be initiated and led by the CEO.” SAP argued in early August that agent sprawl has made AI governance a board-level concern. What I have not found in that work is a standing owner for what separately approved systems do to each other.

A harder objection is that the industry has known about hidden technical debt for a decade and naming it did not fix it. The paper’s prescription was a change in team culture, addressed to the people building one system. Complexity debt accumulates between teams, so no team can pay it down or see both sides of an interface it did not know existed.

Ownership is the link an organization can reach. Naming the person accountable for the composite comes first, and the measurement follows, because someone then needs it to do their job. That person is the chief executive, or an officer the chief executive empowers. Each team is optimizing a local metric, and every one of those metrics can read green while the composite fails. The cost surfaces where the chief executive is accountable, which makes them the first person for whom paying it down is rational.

They build structure, because it’s not possible for an organization to sustain human diligence at machine speed. That means instrumenting the organization so its coordination state is visible, constraining interfaces to make whole categories of incoherence hard to express, and partitioning domains to keep a failure in one from spreading. Human judgment then handles the exceptions the structure surfaces.

That capacity is what I mean by coherence: a company’s capacity to see what its autonomous systems are doing, judge whether they are doing it well, and correct them when they are not. It is the signal box, and it is the part of the railway almost nobody has built yet.

The full argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.