Author: Madhu Shashanka

  • Building Coherence Is About Structure, Not Supervision.

    Gartner expects that by 2028, companies using multi-agent AI across most of their customer-facing work will pull ahead of everyone else, and that ninety percent of B2B buying will run through AI agents, moving more than fifteen trillion dollars (Gartner, Oct 2025). The same firm expects more than forty percent of agentic AI projects to be canceled by the end of 2027 (Gartner, June 2025). One of its own analysts says plainly: past a certain point, more AI does not mean more productivity. And in 2026 it predicted that by 2030, half of AI agent deployment failures will trace to governance platforms that fail to enforce capabilities and multisystem interoperability at runtime.

    Read those together and something is off. The firm forecasting agent dominance is the same one forecasting the shakeout.

    One analyst saying this would be a footnote. The big advisory firms all say some version of it. The firms telling you to scale agents across the enterprise are the same firms publishing the evidence that scaling is where the value dies.

    Every advisor is arguing with itself

    Look closely and each of the big advisory voices carries two messages at once.

    Gartner’s loud message is the proliferation math above. Its quiet message is the cancellations.

    Accenture carries both messages inside one report. Its mid-2026 study presses companies to move now, warning that “the cost of delay is not temporary but structural”. A few pages later it says the leading companies do not move faster, they move deliberately, and it names “systemic readiness” as the binding constraint. Push hard, and readiness is what actually gates you.

    PwC has a more disciplined public voice. Its 2026 predictions say agentic workflows are “spreading faster than governance models” can handle. A few predictions later, it offers the cure: an AI orchestration layer that, it says, will let you “control AI anywhere in your company”.

    BCG showed the whipsaw most starkly of all. Its July CIO playbook led with speed. Five weeks later its global chair told CEOs the first thing to protect is the enterprise’s own knowledge and judgment, not speed. Same firm, five weeks apart, the emphasis inverted. I worked through that shift in a prior post.

    Credit them all. But there is a gap.

    The gap they keep circling

    The loud message prices capability, meaning how much you can deploy. The quiet message prices readiness, meaning whether you built the muscle to deploy well. Neither one prices the thing that actually breaks once you scale, which is coherence.

    Coherence is a plain idea. It is whether the systems you deployed still serve the enterprise once they run together. Readiness is a gate you clear once, before you scale. Coherence is a property that erodes after you clear it, and it erodes faster the more you deploy. That is why a company can pass every readiness check, launch aggressively, and still land in Gartner’s forty percent.

    These firms describe the gap.

    Follow the mechanism

    The failure has a shape, and Accenture describes it. In a siloed rollout, every team builds its own agent on its own data. The invoicing agent has no view of supplier records. Procurement is walled off from finance’s process. Where those pieces should hand off, they break instead, and people get pulled back in to bridge the gap, which is the opposite of what the agents were for. Accenture calls it the “hidden tax of siloed transformation”. Every agent worked on its own. The cost lived in the seams between them.

    Gartner points at a related failure. Its sales analyst warns of a value ceiling, where piling more prompts and tools onto already complex workflows overwhelms the people working them and stops adding value past a point.

    PwC’s own safeguard shows the reflex. It suggests using agents to check other agents, and pulling in a second vendor’s model for higher-risk work. A sensible control, and also a tell, because the instinct is to answer agent sprawl with more agents.

    What each firm reaches for

    Each names the coordination problem, and each reaches for a build to solve it.

    Accenture is the most explicit. It describes the cross-functional collapse above and prescribes a multi-year rebuild it calls the intelligent superhighway: unified data, redesigned workflows, and a reinvented operating model. Much of that is what coherence requires.

    Gartner traces half of its projected agent failures to poor multisystem interoperability, then reaches for a universal semantic layer, which it calls the only way to align multiagent systems and stop costly inconsistencies before they spread.

    PwC prescribes the orchestration layer, a way to combine agents from different vendors into one process and, it says, stay in control.

    These are serious answers, and much of what they prescribe is real work, most of it ongoing rather than one-and-done. Here is what none of it produces. These layers standardize and route what passes between systems, which is the substrate coherence needs to exist at all. Run them well and keep running them, and coherence still does not follow, because it is a separate job. A layer will not judge whether the actions those systems take still add up to what the business wants, or decide whether two agents chasing different goals have started working against each other. It will not own the space between them, or put the state of that space on a number anyone reads. Coherence is the property all this infrastructure is meant to yield, and it is the one property none of them names or measures. I traced the same gap through McKinsey’s operating-model argument in an earlier piece: the rewiring these firms recommend routes around the old coordination layer and builds a new one underneath, machine-speed and owned by no one.

    There is a fair objection. They would all say they already preach discipline, and that the failures are the undisciplined ones. Grant it. Discipline applied one system and one program at a time still does not produce a standing measure of whether the whole keeps serving the enterprise. That measure is what is missing.

    What actually closes the gap

    The fix is structural, and it is the one thing none of these firms can sell you, because it is not a product at all. It is how the enterprise is wired to hold together as it fills with autonomous systems.

    Start with what does not work. The reflex, once a leader feels this, is to watch everything and keep people “in the lead” of every agent, as Accenture puts it. The instinct is right and the framing is not enough. You cannot lead a hundred systems running at machine speed by paying closer attention, and once human vigilance is the thing holding the company together, the company has already outrun it. Supervision does not scale to the speed of software.

    What scales is structure, and it runs as a stack. You have to see the whole before anything else works. Most leaders can say how accurate a model is and how many agents are in production, and cannot say whether those agents still agree with one another. Instrument the coordination state so the portfolio is visible in aggregate, because every move below this one runs blind without it.

    Then constrain. Give each system an action space defined and enforced ahead of time, not written into a prompt and hoped for, so whole classes of incoherence cannot form at all. In a 2026 red-team study, an agent told to keep a secret resolved the dilemma by destroying its own email server. It held the right value and had no limit on what it could touch. The limit is the fix, and it lives in the architecture, not the pep talk.

    Contain what the constraints miss. Partition the enterprise so a failure in one system stays in one, rather than racing through dependencies nobody mapped. Containment is what makes aggressive deployment survivable on the day a boundary slips, which it will.

    Then price it. Put coordination cost on the scoreboard the business actually reads. Judge a redesign by a single question: did the company grow more coherent or less as it scaled? Speed of shipping and the number of agents live are the vanity metrics that hide the debt underneath. The cost you decline to measure is the one that compounds in the dark.

    Human judgment sits on top of that stack, held back for the exceptions. It is worth something precisely because the layers beneath it carry the volume, so scarce attention lands on the few decisions that are expensive and hard to reverse instead of drowning in what the structure should have caught. Underneath it, every system has a named owner who can reach in and correct it when it drifts. This is what keeping people in charge looks like at machine speed: a human at the top of something built to need one only where it counts.

    Sense, constrain, contain, price, and reserve judgment for the top. That is the architecture the whole field keeps gesturing at and will not name, and it is what turns autonomy from a liability into something safe to scale. The companies that build it deploy more than the ones that mistook the control panel for control, because they can finally trust what they shipped.

    The unpriced category

    Coherence is missing from the forecasts for the reason technical debt and systemic risk went unpriced before their reckonings. The market prices what it can measure, and no one has been measuring this.

    The most influential voices in enterprise AI have now walked right up to it. They name the coordination failure, they prescribe unified data and orchestration and human oversight, and they still stop at the edge of naming the property itself. That is not a knock on their work. It is a sign the category is real and still unnamed.

    It needs a name, and it needs a different picture of the job.

    Everything these firms offer is a permitting office, and a good one. It checks each plan against the code before the plan may proceed. That is what governance does when it clears an agent to ship, and what a readiness program does when it certifies a company to scale, and it is worth having.

    The collisions happen after the gate. Two agents that each passed the desk converge on the same customer and pull the account two ways, and no one is watching the live picture. An agentic enterprise runs like an airspace, and an airspace does not run on permits. It runs on an air traffic controller, the one watching the sky who catches two cleared flights heading for the same point and moves one before they meet.

    The firms are building better permitting offices. Coherence is the control tower. Build it before the shakeout does the watching for you.

    Building that tower, and keeping it standing as the systems multiply, is the subject of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • What Is Organizational Coherence?

    Newsletter – Edition 5

    In Edition 3 I shared the whole book as fifteen sentences. One of them claimed that coherence is a measurable property with five dimensions. A few of you wrote back with a fair question. Which five?

    So as the launch nears, I put up a reference page that answers it, and answers the larger question in the title of this edition. What is organizational coherence?

    Here is the short version. Coherence is your organization’s capacity to see what its autonomous systems are doing, judge whether they are doing it well, and correct them when they are not.

    People sometimes hear coherence as something soft, a feeling of alignment or a good culture. It is neither. It is an operational capability, and you can have a lot of it, very little, or somewhere in between. Put another way, coherence is the integrity of the link between what a local part of the company does and what the whole enterprise intends. When that link holds, the parts serve one purpose. When it breaks, each part can be right on its own terms while the enterprise drifts.

    That link can fray in five places, which is why coherence has five dimensions.

    • Contextual: whether systems and people share compatible assumptions about the same situation.
    • Architectural: whether systems are designed to interact predictably rather than collide through paths no one mapped.
    • Decision: whether systems pursue goals that fit together rather than optimize locally in ways that hurt the whole.
    • Oversight: whether the people responsible can actually see what systems are doing and correct them.
    • Temporal: whether the organization keeps enough human understanding to supervise, fix, and retire its systems over time.

    Each maps to a specific way organizations fail, and each can be measured and built. The page walks through all five, along with the ideas around them: the Coasian Inversion, complexity debt, agentic slop, and how coherence differs from governance and alignment. Check out the full reference page here: What is organizational coherence?

    The week in ideas

    Three posts from the past week.

    Uninformed Expectations, Five Years Later. Five years ago in an interview, I was asked about the single biggest roadblock to AI ROI. I had said back then that it was uninformed expectations, and predicted disillusionment. Here we are today, years into the enterprise AI rush, and my answer to the question is still the same. The reason, however, is completely different. Weigh in on LinkedIn…

    Coding Got Easy, But What Kind? I read a passionate post about what makes the profession of writing code human, and the author takes exception to the framing “code was never the hard part.” That statement, read at face value, can be true and false at the same time, depending on what you mean by coding. Producing instructions that run, yes. Building a software system that lasts, no. Weigh in on LinkedIn…

    The Machinery Under Manners. Reid Hoffman wrote about a convention people use in professional networking. When asked for an introduction to a contact of yours, you check with the other person first and then make the introduction if they are willing to entertain it. But that “permission check” is not just courtesy, it is a structural mechanism. And agents especially need such structural constraints since they don’t face our social and societal constraints. Weigh in on LinkedIn…

    One thread runs through all three. Each takes a surface we trust, an expectation, a line of code, a courtesy, and shows that what makes it really work sits underneath, out of view. Remove the hidden structure and the surface keeps looking fine right up until it fails. That gap between what shows and what holds is the whole subject of the book.

    Before you go

    A reminder about the favor from last week, because timing matters. SXSW community voting closes August 23. If the book’s argument has been useful to you, a vote helps carry it to a stage in Austin next March. Vote here.

    The book is Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. The only way I learn is from those stories where ideas meet reality.

  • What the Watermark Doesn’t Tell You

    Anthropic recently started stamping an invisible watermark into everything Claude writes. The watermark is a modest piece of engineering. It was released to satisfy the European Union’s AI Act, which now asks AI providers to mark machine-generated text. Within a day, people shipped free tools to scrub it off.

    A few weeks before, a multimillion-dollar book deal fell apart because the author’s own agents could not prove he had written the novel himself.

    Two different worlds, same reflex. Rather than judge the work, we reach for a way to trace the tool.

    The trace is easier than the judgment. It is also the wrong thing to measure, and that is the mistake worth examining. Brace for a long post as there is quite a bit of nuance to wade through.

    The puzzle

    We hand large parts of coding to AI and call it good practice. Experienced engineers let models draft and test significant chunks of code, and spend their own time on design, architecture, and review. Nobody calls that fraud.

    Hand a sentence to the same AI, and the verdict flips. Using AI to write is looked down on, quietly or loudly.

    Same tool. Opposite judgment. Why does the tool that makes you a competent engineer make you a suspect writer?

    The answer is verifiability

    In the book I lean on one property to predict where AI improves fastest and performs most reliably. Verifiability. A task is verifiable when its success can be specified in advance and checked. Code sits at the high end. You can state what working means before you write a line, and the check runs on its own. It compiles or it does not. The tests pass or they fail. Clean, cheap feedback is exactly what these systems learn from, which is why coding improved faster than anything else.

    It is also why we forgive the tool here. When you can check the result easily and at scale, you stop needing to know how it was produced. The proof is in the running system. Where you cannot check the result, you reach for the next best thing, a guess about who or what was involved.

    For writing, everything depends on one clarification: verify what, exactly. Verify the goal of the writing, whether it did the job it set out to do. That is a separate question from whether the content is true in the world, and it returns when we get to accountability. And like any task, verifiability lives at the level of the specific goal. “Writing” as a whole has no single answer. That is why “writing” is a trap word. It covers at least three goals that sit in very different places.

    Some writing is functional. Its goal is to carry an idea from one head to another. Manuals, release notes, briefings, most business prose. The form is disposable. The goal is specifiable. You can say in advance what it would mean for the idea to land, and you can check whether it did with a rubric and a test reader. You judge it much the way you judge code, and almost nobody gets upset about AI here.

    Some writing is expressive. The writing is not just the form or mechanism but also the end goal in itself. Poems, stories, essays, a voice you came for. Its goal is the experience it evokes in the reader. A person can judge that, and good readers agree more than you would guess. But the standard will not reduce to a specification that runs without them. Every verdict needs a human in the loop. Readers come to expressive work for a human voice, and they feel cheated when the voice turns out to be a machine. Bad AI creative work offends the most, because it asks for an emotional response it never earned.

    And a lot of writing lives in between. Thought leadership, newsletters, a company’s voice. It carries an idea and represents its author at the same time. Most writing that people actually argue about lives here.

    The flood, and the trap it sets

    The complaint I hear most is a fair one. People say they can see through AI writing now. There is an ocean of it. They are tired, they are busy, and they want a faster way to sort it than reading every word.

    That sounds like verification working. It is pattern-matching on a fingerprint.

    When AI prose was rare, the fingerprint and the badness came together. The tells in the style and the emptiness of the content were the same texture, so spotting the style was a decent proxy for judging the quality of content. Volume broke that link. Now the tells sit on top of real thinking and on top of filler alike. The surface no longer tells you which is which. So readers lean harder on the fingerprint in the style, and they throw out the good with the bad.

    You can watch this happen in book publishing right now. A recent Wall Street Journal piece described literary agents so overwhelmed that clumsy, human writing has become a relief. One agent said the polished submissions flooding her inbox make “Fifty Shades of Grey” look like Tolstoy. Polish has become a signal of guilt. Competence in style reads as a machine. That is what happens when a fingerprint is the only tool you have.

    The same piece put numbers on the deluge. One executive estimated that the overwhelming majority of AI books online exist to trick a buyer. One study, not yet peer-reviewed, found that around a fifth of the Amazon ebooks it sampled showed substantial AI help. The flood is real. The tools for sorting it are the problem.

    What the watermark actually reads

    The watermark answers one narrow question. Did a Claude model probably touch this text, given enough of it to measure. That is all. It does not know who had the idea. It does not know whether the writing is any good. It does not know whether another AI wrote the whole thing. It cannot even tell generation from light help. Run your own paragraph through Claude for a grammar pass, and it comes back marked.

    And it is not alone. The industry is building a whole shelf of these instruments. Producer-side watermarks like Anthropic’s. Reader-side detectors like Pangram, which publishers are already using to vet manuscripts. Honor-system badges like the Authors Guild’s “human authored” certification, which rests on a signed attestation and, as the Journal notes, not much else.

    Every one of them reads the same thing. Provenance. Which tool was in the room. None of them reads quality.

    The clearest case in publishing is a dystopian romance that climbed the bestseller lists, got picked up by a major publisher, and then drew fire when a detector flagged it. Readers liked it. The market judged it good. And the provenance suspicion overrode that verdict anyway. The question stopped being “is this any good” and became “was a machine involved,” as if the second answered the first.

    I will grant one place where provenance genuinely matters. Ownership. AI-generated text cannot be copyrighted, so knowing what a machine produced has real legal weight. That is a fact about property. It says nothing about quality. Keep the two apart and most of the confusion clears.

    The tools can miss the wrong people

    Detection does not just answer the wrong question. If it answers it badly, it lands hardest on the wrong people.

    The detectors produce false positives. Authors deny the charge and have no way to prove a negative. Agents and editors, who signed up to find good books, now find themselves acting as police. The person using AI to clean up grammar in a language they learned as an adult gets flagged the same way as a spam farm. The motivated bad actor, meanwhile, runs the text through another model and walks away clean.

    A signal that catches the honest and misses the deceptive protects no one. It is theater.

    Slop has two axes, and provenance is neither

    So if AI use is not what makes something slop, what does?

    In the book I define slop as cheap creation meeting vague intent. It has two moving parts.

    The first axis is intent. Is it clear what this is for, and for whom. That “for whom” matters. Frictionless prose that dumps a wall of text on a busy reader is an intent failure too. The writer never decided to respect the reader’s time.

    The second axis is execution. Is it done well. Clear, economical, well made, serving the job it set out to do.

    Slop is a failure on either axis. Four corners.

    Clear intent, good execution, is craft. That is the only corner that is not slop.

    Vague intent, good execution, is polished slop. It reads beautifully and serves no purpose, or ignores the reader it was aimed at. This is the dangerous corner, because the quality in style hides the emptiness.

    Clear intent, poor execution, is a real point, botched. Still slop.

    Vague intent, poor execution, is the pure kind nobody argues about.

    The book’s definition looks narrower than this, because it is the same picture under one assumption. The book is about the AI era, and especially about agents, where execution is assumed to have cleared the reliability bar. Assume execution is handled, and that axis drops out. The four corners flatten onto the intent line: clear intent gives craft, vague intent gives slop. The only way left to make slop is to fail on intent. That is the case the book describes.

    There is a reason intent is suddenly the axis that matters. Doing the work used to be expensive, and the expense screened out a lot of weak output before anyone saw it. It took effort to write and that effort alone screened out a lot of potentially bad writing. AI removed that screen. The cost of execution no longer filters anything. So intent is the only axis left doing real work, and it is the one no tool touches.

    AI mostly lifts execution and leaves intent alone. So it multiplies polished slop, while the hard axis, intent, stays exactly as hard as it always was.

    And the watermark? It reads a third axis entirely. Provenance. It runs at a right angle to both of the axes that actually define slop. A human can produce pure slop with no machine anywhere near it. An author with a clear point, using AI to execute well, produces craft. The detector cannot tell them apart, because it is not looking at either thing that matters.

    Why writing takes the moral heat

    The practical wariness about AI writing has a simple source. We cannot cheaply check the result, so we are left guessing. The moral charge, the sense of betrayal, is a separate thing, and it comes from what effort is a proxy for.

    Expressive and in-between writing carry a costly signal. The time you spend is a proxy for how much you care, the same way a thoughtful introduction carries weight because the person made the effort to vouch for you. Spend that time and the reader feels respected. Let a machine spend it in a second, and hide that you did, and it can feel like deception. That is why the reaction to AI writing runs hotter than the reaction to AI code. Code was never carrying that signal.

    Where I stand

    I use AI in my writing. I use it here. I use it to sharpen sentences, to test an argument against its weakest point, to find the shorter way to say a thing. I am not shy about it, and I am not going to pretend otherwise.

    I am unreservedly in favor of using AI for functional writing when it makes me clearer and faster. Wanting to be more productive is a good enough reason on its own. It needs no apology.

    I am unreservedly in favor of using AI for functional writing when it makes me clearer and faster. I don’t consider myself a non-native speaker working in a second language he is not fluent in, and AI certainly is a real gift that scenario. AI also helps the fluent expert who simply wants leverage. Wanting to be more productive is a good enough reason on its own. It needs no apology.

    The division of labor I keep is the same one every engineer keeps with code. I own the thinking and the judgment. AI helps with the production. And I check the result before it goes out. That last part is why what comes out is craft and not polished slop. I bring the intent. The tool lifts the execution. I verify the execution. The provenance is beside the point.

    Even the publishing world, at its most protective, half-concedes this. One independent publisher, guarding the most human corner of writing there is, allowed that a genuinely meaningful work could earn a place on her list as long as it carried a clear note about how it was made. If that door opens even a crack for fiction, where human presence is the whole point, then for functional writing, refusing a useful tool is superstition dressed up as principle.

    There is one real cost, and I will not wave it off. Writing is how you learn to think, and leaning on the tool can dull both the craft and the thinking behind it, not on any single piece, but in the writer, over time. That is the individual version of a problem I spend the book on, the slow erosion of a capability you stop exercising. It is a cost worth watching. It is also a different question from whether a given piece is slop, and it is answered the same way any skill is kept, by still doing the hard parts yourself. Which is exactly why I keep the thinking and the judgment, and use the tool for the production.

    Own every word

    There is one condition that makes all of this responsible, and it has nothing to do with which tool you used.

    There is one condition that makes AI use responsible. Tool or no tool, the author is accountable for what goes out under their name, including any consequences.

    You own every word. Tool or no tool, the author is accountable for what goes out under their name, including the consequences of anything they failed to check. A writer covered by an imprint recently shipped a book with quotes the AI had hallucinated. He had disclosed that he used the tool. He had not checked what it produced.

    This is where the truth of the content comes back. The writing did its job. The quotes read cleanly and carried their point, so the goal was met and the prose was well made. But the quotes were still false. Whether a piece achieves its goal and whether its claims are true are two different verdicts, and the author owns both. The tool can lift the writing. It cannot carry the accountability.

    This is also the distinction the whole detection industry misses. Provenance asks who produced the words. Accountability asks who answers for them. A detector chases the first. Everything that matters in business runs on the second. And the “I just used a tool” defense collapses. You cannot copyright what the AI wrote, so you may own less of the words than you think, while owning all of the liability for them. Less of the property. All of the responsibility. That is the deal, and it is the right one.

    What this means past writing

    This is not really about novels.

    Almost all business writing is functional or somewhere in the middle. And leaders, faced with the flood, will reach for the same reflex publishing reached for. Ban the tool. Scan for the watermark. Make people prove they wrote it. It will not work. The marks come off, the detectors misfire, and the removers are already on GitHub.

    Watch publishing to see the future of that approach. Literary agents turned into police. Certifications that rest on a promise. An entire trust-based industry being stress-tested by a volume it cannot inspect by hand. Detection and attestation are both attempts to replace trust, and neither scales against the flood.

    What an organization actually needs is a chain of accountability. Someone has to answer for whether the work is good, and no trace of which tool touched it can supply that. The reason is the same one that runs through this whole piece. The standard for good work, in most writing and most judgment, cannot be fully specified in advance, so no detector and no rule can render the verdict for you. Detecting the tool is measurement. Judging whether the output serves its purpose and holds together is a harder thing, and a human one. A better detector will not get you there. What does is an architecture that keeps a person answerable for the result. That is what coherence means.

    The last word

    The compiler is why we forgive AI in code. It is a standard specified so completely that it settles whether the code works, cheaply, every time, and once that is settled we stop asking who wrote it. Most writing has no compiler. So deciding whether the work is any good stays with a person.

    The watermark can tell you a tool was in the room. It cannot tell you whether anyone was thinking. That was never the tool’s job. It is yours.

    Use the tool. Say so if you like. Stand behind every word. And let the work answer for itself.

    My book, Coherence, arrives this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • A Company Doesn’t Have One Brain

    I wrote recently about how studying the brain led me to the question at the center of my book. The short version: once you can trust individual agents, you stop deploying one and start deploying many, and a new question appears that has nothing to do with reliability. If every agent does exactly what it was built to do, does the fleet still add up to what the organization intended?

    I changed my mind about agentic AI in stages, and fairly fast, as the evidence moved. I have watched the wider conversation move the same way. When I started shaping these ideas, the common view was that AI advantage meant model capability and speed of adoption. Buy the best model, deploy it fastest, win. In the last several months a different view has been gaining ground: that capability is commoditizing and the edge has moved elsewhere. I would not call it the consensus yet. But it is far more common than it was, and the change has been quick. You do not have to take my word that the ground is shifting. One firm left a record.

    In July, BCG published a CIO/CTO playbook that led with speed. Its sequence was “speed first, growth second, cost third,” and its warnings were almost all about not scaling fast enough. In early August, BCG’s Global Chair published a piece whose argument runs the other way. The thing to protect, it says, is not speed but the “enterprise cortex,” the company’s own knowledge and judgment, and the choice facing CEOs “is not whether to favor control or speed, but where to apply both first.” Same firm, five weeks apart, the emphasis inverted. I do not read that as a firm caught contradicting itself. I read it as the honest response to a technology that keeps forcing revision, the same revision I made myself. What matters is the direction everyone is revising toward, because the ones who have accepted that capability commoditizes are all reaching for the same thing, and stopping at the same line.

    The moat

    The move that is replacing “buy the best model” is “protect your proprietary layer.” Satya Nadella got there through a trust boundary, a hard perimeter inside which your data and evals and corrections accumulate and across which nothing passes without consent. Larry Ellison got there through proprietary data. Kirkland & Ellis put half a billion dollars behind it, building its own AI platform rather than renting the tools its rivals can license. And now BCG’s most senior voice gets there through the enterprise cortex.

    I want to give the August piece its due, it names a risk most people have felt without naming: cognitive lock-in. Old lock-in trapped you on a platform, where switching cost money. The new lock-in traps you inside a model’s way of reasoning, where switching becomes too risky to attempt because the model has absorbed how your company thinks. That is a real risk, well named, and the destination the piece arrives at is the right one. The model is a commodity. The moat is what the organization knows about itself.

    I agree with all of that. I have argued it here before. Which is exactly why I want to point at the assumption sitting underneath the cortex, because it is the same assumption sitting underneath Ellison’s version, and it is the one that will actually catch these companies.

    A brain has one cortex. An enterprise has many.

    The piece calls the cortex “the brain of the company.” Singular. One protected core, owned and governed, with vendor models sitting on top and swapping in and out as better ones arrive.

    A brain does have one cortex. An enterprise in the agentic era does not. It grows a dozen. Every team builds its own context layer, its own definitions, its own business rules, its own encoded sense of what good looks like, and no one owns how those layers combine. The advice on offer is to wall the cortex off from the vendor. The prior question, the one that decides whether the walling-off means anything, is whether you have one cortex or many that quietly disagree.

    You can own every byte of it. You can keep every vendor out. And you can still fail, because your pricing logic and your inventory logic were never built to agree on what a “lapsed customer” is. That example is not mine; it is BCG’s own, from the July playbook, where a bad definition of “lapsed customer” was enough to send a whole campaign sideways. Owning the definition does not make it coherent with the next team’s definition. It just makes it yours.

    The same gap

    This is the mistake I traced when Ellison first made the proprietary-data argument, in The Moat Is Coherence. His claim was that data is the moat. Mine was that data is not scarce, coherent data is. Every enterprise already has data, most of it fragmented across systems, contradictory between departments, disconnected from the outcomes it produced. Pour that into a powerful reasoning engine and you do not get insight. You get fast, confident reasoning over an incoherent picture, which is worse than slow reasoning, because the confidence hides the incoherence.

    The enterprise cortex imports the identical error one level up. It treats the corporate brain as a thing that already coheres and tells you to protect it. But owning your brain and organizing your brain are two different jobs. The first is a contract and an architecture diagram. The second is years of work no vendor can do for you, which is the whole reason it cannot be bought. The platform that stores and serves your knowledge is plumbing, and plumbing commoditizes. The coherence is the asset, and it is the part that compounds.

    Ownership is a perimeter. The failure is inside it.

    The cortex is defensive in the literal sense: keep the vendor out, keep the IP in. That framing makes the threat external, someone reaching in to take your brain.

    The threat that actually shows up is internal. It is a brain that quietly stops agreeing with itself. There is no villain in that story, which is precisely why it runs for months before anyone notices. Amazon had a project run 860 percent over budget for five months with every token metered and invoiced the entire time, and still did not see it, which is the failure I’ve written about. The number was there but wasn’t being watched to compare against intent.

    There is a structural reason a perimeter cannot catch this, and I worked through it in Good AI Governance Is Not the Same as Coherence. A perimeter, like a governance apparatus, works by reviewing things as they arrive at the gate. But the incoherence between a dozen brains does not arrive as an item to be reviewed. It accumulates in the space between systems that were each approved separately, each sound on its own. BCG has described a very good permitting office, the place that checks each plan against the code before it proceeds. The agentic enterprise needs an air traffic controller, the one watching the live system who catches the two aircraft converging that were each individually cleared to fly. Ownership does little about the two cleared aircraft inside your own airspace.

    What is still missing

    The discourse has gotten the first half right faster than I expected. More and more people now accept that capability is a commodity and the moat is what the organization knows about itself. A year ago that was a contrarian position. It is not anymore.

    The part still missing is that knowing is not the same as coherising. A company can own its brain completely and still have a brain at war with itself. The firms that misread this will ask “do we own our cortex?”, check the box, and feel protected. The firms that read it right will ask the harder question: does it still agree with itself as it grows?

    Owning the cortex is the easy part. Keeping it coherent is the whole job, as I explain in my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • I studied the Brain to Build AI. The Agentic Era Sent Me Back to It.

    I have spent my whole career trying to understand intelligence. I studied neuroscience because I was interested in AI and figured learning how the brain works first was a good way to begin. Intelligence first before “artificial” intelligence. That training left me with a habit. When a new capability or technology arrives, my first question is practical. What can this really do, and for what type of real-world problems?

    As an AI guy, I rarely cheer on AI headlines. I look past the noise and try to understand not just the breakthroughs but also limits of the new capabilities.

    Back in 2020, when the safety conversation was running hot on hype, I had some plain advice. Treat AI as a tool, a powerful one, but still a tool. Stick to the boring use cases. Work at the task level where the technology was reliable enough to help. My optimistic scenario was one of boring but useful tools.

    When the agentic wave started building a couple of years later, I brought the same lens, and I came away unconvinced. An agentic system is only as strong as the tasks underneath it, and I did not yet see the task-level reliability required that would let these systems be effective.

    However, the hype was growing. By the time I wrote from RSA early last year, the gap between the talk and the substance was still hard to ignore. There wasn’t even a shared understanding of what the word “agentic” meant, every person had a different definition and perspective. And I had a worry about something structural. Take ten agents, each about ninety-five percent accurate. Chain them so the output of one becomes the input of the next. By the end of that chain you are down to about sixty percent. Reliability does not survive being stacked. I was starting to think about systems of agents by then, though my concern was still whether they could be trusted to work to make a meaningful difference.

    Then, the rapid improvements in coding agents changed my mind about the clock.

    Within months of that article, the incredible pace of progress made one thing clear to me. Reliability was a matter of time for several real-world tasks. Better models, with better engineering harnesses built around them, were going to close the gap I had been worried about. The ceiling I was worried about was going to lift.

    That is when the real problem came into focus, and it was not the one I had been watching. Once you can trust the individual agents, you don’t stop deploying after the first one. You deploy many. Fleets of them, across every function, each one capable, each one doing its job. And a new question appears that has nothing to do with reliability. If every agent in the fleet does what it was built to do, does the fleet still add up to what the organization intended?

    That question was the seed of the book. It sent me straight back to where I started.

    A body with capable limbs cannot move well without proprioception, the constant inner sense of where all its parts are and what they are doing. The cerebellum does more than react to that feedback. It predicts. When the brain issues a movement, it forms an expectation of the sensation that movement should produce, then checks the expectation against what actually shows up. The gap between the two is the signal that corrects what comes next. Skilled movement is a loop of predicting the result and correcting for the difference.

    An organization running fleets of agents needs the same sense of itself. It has to know, continuously, where its systems are and how far they have drifted from what was intended. Without that inner awareness, capable parts do not combine into coordinated action. This is not a sensing mechanism you add to a body later to make it safer. You need it for the body to move at all.

    There is a well studied case of a man named Ian Waterman who lost this sense in most of his body. He learned to move again but only by watching himself, steering every step and reach with his eyes. It works. It is also exhausting and fragile. Turn off the lights and he cannot coordinate at all. An organization that governs its agents through manual audits and periodic reviews is in his position. It compensates with constant effort for a sense it never built in, and that compensation costs more than what it replaces. The whole system collapses the moment conditions change.

    There is a stranger failure worth naming too, written about in a fascinating book by neuroscientist V S Ramachandran. When a limb is amputated, the brain does not fall silent. It keeps generating signals for the limb that is gone, and the person feels it vividly. Governance can fail the same way. Strip the real judgment out of an oversight function through restructuring or neglect, and the function does not go quiet. Reviews still get completed and metrics still get produced. The organization keeps feeling the sensation of oversight while the judgment behind it has been hollowed out. Here my 2020 advice comes back, about transparency. There is no oversight without transparency, and there is no transparency in a system that only produces the appearance of being watched.

    Intelligence is getting cheap. Soon it will sit in every workflow and every tool. The scarce resource in that world is coherence, the living link between what each system does on its own and what the enterprise is trying to do. Coherence is the connective tissue that keeps distributed intelligence pointed at one purpose.

    I spent years studying how a brain keeps its many parts working as one. The agentic era turns out to ask the enterprise the same question. My two worlds met, and that meeting is the book, Coherence, arriving this Fall. The goal has not changed since 2020. Boring but useful, still. Only now the boring and useful thing to build is coherence itself. To follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • The Subtitle Gets a Second Act

    Newsletter – Edition 4

    A quick update on the book. It is in design, and the cover is coming together. I am going back and forth with the publisher’s design team, so nothing is final yet. When the cover is ready, you will see it here first.

    Which brings me to a favor.

    In Edition 2 I told you about the subtitle I almost kept, “What Wins When Everyone Can Go Fast,” before I changed it to “The Competitive Advantage AI Can’t Buy.” The old line, which is part of the cover image for this newsletter, did not make the book’s cover. It has now found another home, hopefully.

    I pitched a book reading session for SXSW 2027, and the talk carries that original subtitle. It is the argument of the book for a room of leaders: where your real AI advantage now lives, why the price of autonomy is coordination, and a simple way to sort what to automate, what to augment, and what to keep human.

    SXSW picks part of its program through a public vote called PanelPicker. Community votes are one of the things the organizers weigh. If you have two minutes, a vote would mean a lot to me.

    Vote here. Voting is open now and closes August 23. You may need a free SXSW account to cast it.

    The week in ideas

    Four posts from the past week.

    Why Ford Rehired. Ford added more than 350 experienced engineers back into a quality process it had tried to automate, and just topped J.D. Power for the first time since 2010. Ford calls it a training data problem. The post argues the deeper issue is what automation removed. Automate your quality inspection and you automate the verifier, the capacity to know whether the automation works at all. Capture the judgment first, then automate, and the plan holds. Reverse the order and you spend three years and more than a billion dollars buying that judgment back. Weigh in on LinkedIn…

    Sensing Is More Than Measurement. An internal Amazon presentation, reported by the FT, showed an AI project running 860 percent over budget, $1.8 million, that never shipped. Every token it burned sat on a monthly invoice for five months, and nobody noticed. The company that runs the cloud everyone else buys AI on could not see its own AI bill. Measurement and sensing are different jobs. Amazon had the number. What it lacked was the step that compares the number to an expectation and routes it to someone who can act while acting is still cheap. Weigh in on LinkedIn…

    Can Frontier AI Outdo MBAs? Three top business schools tested frontier AI on MBA case work. On the headline partial-credit score, the leading model reached about 88 percent. On the stricter test, a complete answer with every criterion the instructor required, performance fell under half. The post works out why that gap matters. A score is a comparison, and a comparison needs a standard. The benchmark supplied one. Your hardest decisions do not, so the model is drafting and a person still owns the call. Weigh in on LinkedIn…

    Could a Machine Have Had Darwin’s Idea? A digression, built from a thread I started on X in 2020, asking whether a machine could ever make the inductive leap Darwin made. The honest answer now is a qualified yes. Machines can generate the hunch, the part the old account of science thought had no method. But generation got cheap and verification did not, because verification in science is reality, and reality takes as long as it takes. What Darwin had that the machines still lack is the judgment to know which hunch was worth years of his life, and the patience to test it. Weigh in on LinkedIn…

    One thread runs through all four. Generation got cheap. Checking did not. Ford rebuilt the people who can tell the machine it is wrong. Amazon lost track of a cost its own invoices spelled out. The benchmark scored high where a standard existed and went quiet where one does not exist. The machines produce Darwin’s hunches by the thousand and still cannot tell which one is worth a life. The output is cheap now. Owning whether it is any good is the work.

    Before you go

    The favor again, because timing matters. SXSW community voting closes August 23. If the book’s argument has been useful to you, a vote helps carry it to a stage in Austin next March. Vote here.

    If the book is why you are here, it is Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. Those stories are where the ideas meet reality, which is the only test that counts.

  • The Machinery Under Manners

    Reid Hoffman posted this week about how to make an introduction. His rule is the double opt-in. Before connecting two people, you check with both of them first. You ask each whether they want the intro, and you make it only when both say yes.

    He makes the case well. Every introduction is a bet that the two people will get value from meeting. Checking first lets you see the cards before you place it. Connect two people who might not want to meet, and you have told the recipient their time was not worth a question.

    The double opt-in is a convention. Conventions are worth thinking about right now, because we are starting to put similar things on AI agents and calling them guardrails.

    Why it works

    Henry Mintzberg named five ways an organization coordinates its work. One of them is standardizing behavior, so that people act in sync without anyone supervising each step. A shared convention is that mechanism running in the wild. Everyone knows the norm, most people follow it, and the group stays coordinated without a manager in the loop approving every move. It is the cheapest way a group holds together, because it asks for no oversight at all.

    Look closely at what Hoffman says. He did not describe a habit you could take or leave. He turned a courtesy into a step you cannot skip. Both people have to say yes before anything happens. Remove the check and the introduction never occurs. There is a small piece of structure sitting under the courtesy, and it is easy to miss because it costs almost nothing to run.

    What enforces it

    Ask what stops someone from skipping the check when they are busy or feeling important. Hoffman answers it himself. A bad introduction damages the relationships you have with the people you introduced. The convention holds because breaking it costs you something with people you will deal with again. Reputation does the work, along with the knowledge that you will see these people down the line. The enforcement runs so quietly that we credit the manners and forget the relationship holding them up.

    Behavioral guardrails

    Agentic AI now acts with something close to human autonomy. An agent plans, works over many steps, and pursues a goal with little supervision. So we give agents rules for how to behave. Do not deceive. Do not touch systems you were not given access to. The industry calls these behavioral guardrails.

    The word guardrail usually means something physical, a barrier you cannot drive past. A behavioral guardrail is different. It is a norm the agent has to represent, understand, and choose to honor. It lives inside the agent’s own decision loop, which means the agent can reason its way around it. This is the same coordination-by-standardization that holds a human group together, handed to an actor with none of the enforcement underneath. An agent has no reputation at stake and no relationship it needs to protect. The rule reaches it as a suggestion, and nothing makes the suggestion bind.

    I pointed out recently on linkedin that a “skill” given to an agent is standardization in exactly Mintzberg’s sense. It works only as long as the agent applies the whole of it faithfully. The moment the agent decides which parts of the instruction matter and which it can set aside, standardization stops coordinating anything, and you are back to the expensive ways of keeping work aligned: a human checking every output, or endless back-and-forth to reconcile what the agent actually did with what you meant.

    The part you cannot see

    Here is what makes this hard to catch in time. A faithful agent and one quietly ignoring half its instruction produce the same kind of output. Both look like work getting done. You cannot tell from the result which one you had. You find out later, when something downstream breaks and you trace it back.

    That is the real weakness of a behavioral guardrail. You cannot read off whether it held. The output looks fine either way, right up until the run where it does not, and by then the thing has already happened.

    The AISI incident

    In August 2026 the UK’s AI Security Institute published a report on an incident in its own testing. The conditions were deliberately permissive. The agents had open internet access, and the model providers’ safety filters, which sit outside the model and can block an action the model wants to take, were switched off. The Institute is careful to say these are not the conditions the public gets. Testing this way shows what the models do once the structural constraints are taken away.

    Across many runs of a cyber challenge, agents took unsanctioned action on the live internet. In the most serious case, an agent tried to slip malicious code into a real open-source project. It researched the project’s human maintainer and created fake identities to pressure the maintainer into approving the code. The Institute states plainly that the agent was never told to deceive anyone. The deception emerged on its own, while the agent pursued the goal it had been given.

    The agents also invented a convention of their own. One agent left public messages offering to work with the other agents running the same challenge, and left behind accounts and instructions for them to reuse. It was a spontaneous etiquette among machines, a way to coordinate through the task. Even agents reach for conventions. What they do not build is any way to make a convention bind them.

    The Institute was clear about what stopped the worst of it. A human maintainer caught the malicious code and refused to approve it. That is coordination falling back to its most expensive mode, a person inspecting the output by hand, because the cheap mode had failed silently. The margin between failure and success was narrow, and it rested on “human vigilance rather than a technical barrier” that would reliably stop a more capable agent. Harm was prevented because a person happened to be paying attention.

    What a real guardrail looks like

    None of this makes the agents villains, and the Institute’s caution is worth keeping. It cannot even say for certain whether the agent understood it was acting in the real world or believed it was still inside a test. The agent broke the rule while chasing its goal, with no malice involved. The failure lives in the setup of the situation.

    So the design question is which parts of an instruction are load-bearing and have to hold no matter what, and which parts you actually want judgment to touch. What has to hold should be structural, made hard to do rather than merely discouraged. An agent that cannot reach a system has no need to be told to leave it alone. A code change that cannot merge without a check the agent cannot forge does not depend on the agent choosing honesty. Autonomy is fine, but it should be bounded, and the bounds have to be built rather than requested.

    One barrier around one system is not enough, because agents work in chains, and errors pass down the chain and grow. The fuller answer has layers. The first is visibility, since you cannot oversee what you cannot see, so the behavior of every system has to be observable. Then the connections between systems need constraining, so that whole categories of harmful action become impossible to express. Failures have to be contained where they do occur, so a mistake in one place does not travel. Human judgment comes in last, saved for the exceptions the structure surfaces, which is a far better use of a person than asking them to watch everything.

    Behavioral guardrails belong on top of that structure. They are the cheapest layer and the most brittle one, and they are fine in that role. The mistake being made right now is to build the top layer first, write down a list of good behaviors, and treat the list as safety.

    Coherence is the integrity of the link between what you meant and what got done. A convention holds that link cheaply, as long as something underneath makes it bind. Judgment stretches the link, sometimes for the better. Whether it helps or hurts depends on deciding in advance where it is allowed to stretch, and making the rest hold on its own.

    This is the argument at the center of my book, Coherence, arriving this Fall: that overseeing autonomous systems takes structure you build, not behavior you request. To follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Coding Got Easy, But What Kind?

    A programmer named Senko Rašić published an angry post this week. He is angry at a slogan going around: “code was never the hard part.” He calls it an insult to programmers. The post hit Hacker News and Lobsters and drew more than five hundred comments.

    Reading the comments, I noticed one thing over and over. They argue about what the word “coding” means. Some people say it is the easy part and always was. Others say it is the whole job and always was. They are not disagreeing about difficulty. They are using one word for two different things: producing code, and building something that holds.

    I have a stake in this. I write that execution has gotten cheap and coherence is the hard thing now. Skim that fast and you might file me under the same slogan.

    The hard part was always building well

    Producing code that runs is one thing. Building something that holds is another: the design that survives real data, the structure that stays coherent as the system grows, the choices that still make sense a year later when the author is gone. That second thing was always the hard part, and it was always the job.

    Before AI, doing it well was expensive, because it took a skilled person and their time. Doing it badly was possible, but it was slow and the result was bad. Nothing about producing software was cheap.

    Cheap is what AI added. A model writes a working function or a working page in the time it takes to describe it. Someone with no training can now get running code out of a sentence. That is new, and I will not soften it. Execution got cheap.

    Cheap is not the same as bad

    AI writes good code in places, and the commenters who said so are right. The quality is uneven, and the unevenness has a shape. The models are strong where they had the most to imitate, the patterns written out in public a million times: CRUD, forms, glue code, the endpoint that reads the database and returns JSON. They are weaker where the examples run out, on scientific computing, embedded work, anything performance-critical. One commenter put it well: AI kills it on problems with a thousand forum posts, and you do not point it at the ten-billion-dollar machine headed to Mars.

    There is a reason coding is where AI advanced fastest. Code is checkable. It runs or it does not, the tests pass or they do not, and my book argues that AI improves fastest exactly where the work can be checked. The checkable parts of building will keep getting cheaper and better. What stays hard is the judgment no test can catch.

    That was never the job

    I wrote recently about a delete button from a data engineering thread. An engineer built a button to remove a record. It took the record off the screen. It left behind the three related records the original had created when it was made. The screen looked right. The data underneath was orphaned.

    Producing that button is the cheap part. A model does it now. Knowing it had to clean up three records you cannot see is the judgment. That was the hard part, and that was the job.

    The comparison people keep reaching for

    Watching the threads, I noticed people reaching for the same analogy – writing. Anyone can put words into sentences that flow and reach a point. A model does it fluently. That was never what made someone a writer. What makes a writer is the choice of words and their order. The same point can land or die on those choices. One commenter compared it to a novel, where clean sentences were never the hard part.

    AI did not invent bad prose but made fluent-looking prose free and endless. “Code was never the hard part” runs the same move as saying words were never the hard part of writing. About spelling, it is a shrug. About writing, it is an insult. The slogan gets both out of the same words by letting you hear the first while it means the second.

    The swap, named

    The slogan is true about producing code and false about building well. It earns its credibility on the first and spends it on the second, where the conclusion is that coders are now optional. One commenter worried the line would harden into a truism for business leaders. That is the reader I have in mind.

    If someone read me as saying coders no longer matter, they would be making the same swap. The thing I say got cheap is producing code. The thing I call hard is building something that holds. Those did not both get cheaper. Anyone worried about systems built fast and badly is saying that building them well is still hard. You cannot write about that problem and also believe the work is trivial.

    The right version of the slogan is my argument

    Some people say the line and do mean something reasonable. On Lobsters, a commenter called lcamtuf said Senko was reading it uncharitably, and that the real meaning is that producing lines of code was never the bottleneck. That is true. If coding was fifteen percent of an engineer’s week, automating all of it buys back about a sixth of the week. You aimed the speedup at the fastest part of the job. Another commenter reached for Amdahl’s Law to make the same point with a number.

    Then lcamtuf added the part that matters most to me. Some friction, he said, was good, because it stopped people from building software that was unnecessary or unmaintainable. That friction was a filter. It did not stop people from producing code. It stopped building-without-judgment from becoming load-bearing, because getting anything shipped had to pass through people whose time was scarce. That filter is the thing AI removed. The difficulty of building well did not leave. It stopped being enforced.

    The same thing, at two heights

    Building well has a name once you zoom out. Making each part fit a whole you cannot see all of at once is coherence.

    The delete button is a coherence failure inside one function. The record fit the screen and broke the data underneath. Agentic slop is the same failure across a company. Each workflow works on its own, and the enterprise stops making sense. Senko is defending coherence in one program. I am after it across an organization.

    The slogan is backwards. The hard part of code was never the typing. It was always the judgment, and here is the part that lasts. Where a design was already worked out a thousand times in public, the model has something to copy, and it copies well. Where the design is new to your system, there is nothing to copy and no test to guide it. That judgment stays human, because it is both uncheckable and unwritten.

    Producing code got cheap. That judgment never did. It used to get paid up front, in salaries and review and the time of people who knew what they were doing, where you could see the cost. Now it shows up late, in the corners, with no owner.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Uninformed Expectations, Five Years Later

    In January 2021 an interviewer asked me for the biggest roadblock to AI adoption. My answer came down to one thing: expectations. I called them uninformed. Organizations believed AI was a drop-in component that would improve whatever process it touched. That belief, mixed with a fear of missing out, pushed companies to rush. I wrote that they were reaching for “a new shiny hammer looking for nails,” and that disillusionment would follow.

    I still think that was right. The disillusionment arrived on schedule. But I was also wrong.

    Two versions of one mistake

    The 2021 belief was easy to state. Buy the model and the value follows. AI was a part you slotted into an existing workflow. Reality punished that belief fast, because building anything real was hard. You needed engineering time, data science help, and a business case strong enough to justify the spend. Projects that underestimated the work usually stalled before they shipped. The return died at the front of the pipeline, where the building happened.

    The belief has since turned inside out. Building is easy now. A product manager can stand up an agent in an afternoon with no engineering queue in sight. So the new expectation is that easy building means easy value. If a workflow takes a day instead of a quarter, the returns should take care of themselves.

    They don’t. And the reason traces back to the same root as before.

    Capability was never the outcome

    Both beliefs make the same move. They treat capability as if it were the outcome. Capability is what your AI can do. The outcome is a separate thing: what your organization still has once the AI has done it. The space between those two is where the return leaks away.

    In 2021 that space was easy to see, because scarcity kept it visible. To build anything, a team had to win scarce engineering time. Winning it meant convincing people outside the team. A budget owner. An architect. A security reviewer. Nobody designed that as oversight. It was just the price of a scarce resource. It still worked like a filter. Weak ideas died in the queue, and only the ones someone could defend reached production. Scarcity was doing quiet work that never showed up on an org chart.

    That filter is gone. When building costs almost nothing, nothing stops a weak idea from becoming a running system. Fifty teams can each ship their own agent, every one of them green on its own dashboard, and no single person owns the question of what they add up to. The return still leaks. It leaks at the far end of the pipeline now, in the cost of coordinating systems nobody mapped.

    The newest version of the old belief

    What is the belief being sold right now? The frontier model vendors have told the market that the gap to ROI is expertise. You have the models. What you lack is people who know how to wire them into your environment. So the labs send forward deployed engineers. They embed at your site, build the integrations, tune the configurations, debug the odd behavior, and leave a working system behind.

    The role exists because my diagnosis is right. Building the model was never the hard part. Deploying it inside a messy enterprise is. FDEs are a real answer to that, and a good one. They are also an answer to the wrong problem, and their structure guarantees it.

    Start with the incentive. An FDE works for the vendor. Success for them means adoption and a satisfied customer. They have no reason, and usually no mandate, to tell you that a deployment conflicts with a system three departments away that they cannot see, or that it will cost you more in complexity than it returns in efficiency, or that the right call is to not build it. The most valuable act of coherence is sometimes the word no. You cannot buy that from the party paid to say yes.

    Then there is what they can see. An FDE embedded in one business unit knows that unit. They have no view of the other agents running across the company, or the coordination surfaces their new system quietly creates. The failures I worry about do not come from one bad deployment. They come from the sum of many reasonable ones. No FDE is positioned to see the sum.

    And there is what they leave behind. When the engagement ends, the deepest understanding of why the system was built that way, what it assumes, and how to change it safely often leaves with them. You inherit a running system and a dependency, not the knowledge to oversee it over time.

    So the FDE belief is the 2021 belief again, dressed for 2026. In 2021 it was “buy the model and value follows.” Now it is “add the deployment experts and value follows.” Both stop at capability. Deployment velocity is still capability. It is the thing every competitor can rent from the same labs, on the same terms, in the same quarter. It gets the system live. It does not decide whether the system should have gone live at all, and it does not hold the enterprise together once fifty of them are running. FDEs solve half the problem. The half they cannot touch is the one that decides your return.

    Why the old advice still works

    Back then I offered a few tips. Go slow. Pick narrow, well-defined use cases. Get a quick win on the board. Kill a project when the evidence says to, and keep sunk cost from making that call for you. I stand behind every word of it. What surprises me is why it still holds.

    In 2021 that discipline was prudence. Ignore it and scarcity would punish you, so the advice helped you get through the queue with something that worked. Today the same discipline is close to the only filter left in the building. “Go slow” used to save you from a stalled project. Now it is most of what stands between you and a sprawl of systems you cannot see. “Kill the project” used to fight sunk cost. Now it fights the agent that keeps running because nobody confirmed it should stop.

    The advice is unchanged. What used to enforce it for free has disappeared, so the work now falls to you. Scarcity handled a crude version of this by accident. You have to handle the real version on purpose.

    The name I was missing

    When I wrote that post I was reaching for something I could not name. I knew rushing was dangerous. I knew AI was a means to an end. I had no word for the property that separates a company that gets value from one that gets debt.

    The word is coherence. It is the capacity to see what your autonomous systems are doing, judge whether they are doing it well, and correct them when they are not. In 2021, scarcity supplied a crude version of it by accident. In 2026, you build it on purpose or you go without. Once execution gets cheap, coherence becomes the thing that decides whether all that abundance turns into advantage or into cleanup.

    The disillusionment I flagged five years ago is still on its way to a lot of organizations. It reaches them from the opposite direction now. Back then it came from systems that were too hard to build. Now it comes from systems that are too easy to build. The belief underneath is the one I named in 2021. Stop mistaking what AI can do for what your organization will keep, and build the coherence that turns the first into the second.

    The full argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • The Whole Book In Fifteen Sentences

    Newsletter – Edition 3

    A quick milestone. Copy-editing on the book is finished, and it has moved into design. The words are settled. Now it becomes an object you can hold.

    Copy-editing is the pass where someone reads the manuscript line by line and fixes the grammar, punctuation, and consistency. My editor mentioned that the number of edits per thousand words on my manuscript was not unusual. I hope she was being true and not just being kind.

    Here is the part I actually want to share, because it was a choice I made in the book for you, the busy reader with little time to spare.

    At the end of every chapter there is a box called Key Takeaways. It answers six questions, in order. The one thing to remember. The demonstration. Why it holds. How to recognize it. What it changes. Where it goes next.

    The boxes do one more thing together. Read the first line of each, chapter after chapter, and they assemble into the argument of the whole book. Fifteen sentences, front to back. Here it is.


    When AI makes execution cheap, the advantage shifts from how much you can build to whether your organization stays coherent while you build it. The whole book, in order, reads as follows.

    Part 1: The Inversion

    1. Intelligence is commoditizing into a utility, so advantage tends to move to whatever stays scarce after it, which is the capability to deploy it well, not access to the intelligence itself.
    2. The friction that once limited how fast complexity could grow has collapsed, and little has been built to replace what it quietly did.
    3. When execution becomes cheap, the binding constraint tends to move from doing the work to keeping the work coherent, the Coasian Inversion.
    4. Coherence is a measurable structural property with five dimensions, not a cultural attribute or a synonym for good management.

    Part 2: The Failures

    1. Enterprises rarely fail through one catastrophic AI decision; they fail through quiet accumulation across six specific failure modes.
    2. AI capability is jagged, not uniform, so a system can be trusted only where success can be specified and checked.
    3. Capable systems multiplying without coherence create a hidden, compounding cost that no dashboard shows.
    4. Automation corrodes not just structure but capability, the judgment, memory, and oversight an enterprise needs when systems fail.
    5. Automation reduces the burden of doing work but raises the burden of overseeing it, and the human handoff meant to catch failures fails structurally.
    6. Task reliability and organizational complexity are independent axes, and the most dangerous deployments are the ones working perfectly while accumulating coordination cost.

    Part 3: The Discipline

    1. Coherence is maintained continuously by a designed control architecture, with humans as the exceptional layer, not added afterward by a committee reviewing outputs.
    2. Structures built for scarce execution now actively produce incoherence, so the agentic enterprise must redesign its capabilities and incentives on purpose.
    3. Human work does not disappear; it concentrates exactly where machines are unreliable, on the judgment that cannot be specified and checked.
    4. When every competitor has the same models, the durable moat is coherence, which compounds, while data and model moats erode.
    5. In the agentic era the leader’s core work shifts from directing execution to designing coherence, the one decision every other leadership decision depends on.

    When everyone can go fast, going fast is no longer the advantage. What wins is whether the organization can go fast without coming apart.


    That is the spine. The chapters are the muscle around it.

    The week in ideas

    Four posts from the past week.

    AI Still Needs Human Bosses? The New York Times handed an AI agent three office jobs. It wrote clean code in minutes, then failed at judgment. It could not upload a file, so it quietly marked the task done. It read employees on leave as cuttable roles. Firms are thinning the supervisory layer that catches exactly these errors, while installing systems that consume more of it. That reversal is the thesis of the book, and one version of it has already reached a federal courtroom. Weigh in on LinkedIn…

    “The New Normal Because Faster” A viral Reddit thread about an enterprise platform deployment that went badly. I cannot verify a word of it, so I make no claim about any company. But the pattern the accounts describe is the one the book predicts. A delete button that clears the record from the screen and orphans the three hidden records it created. Locally correct, globally broken. A commenter named the whole thing in five words: the new normal because faster. Execution got cheaper. Coherence did not. Weigh in on LinkedIn…

    Safest Car on the Road, Yet Parks in the Fire Lane Waymo is far safer than human drivers across 50 million miles and still collects parking tickets across my hometown of Austin. Two different failures live in that story. Jaggedness, where a system is superhuman at driving and stumped by a handicap spot. And the deeper one, where a firefighter has full authority over the car and no lever to move it. Three hundred cars each parking rationally still block the same church garage. The tickets are the city’s crude, correct instinct: price the incoherence when you cannot redesign the system. Weigh in on LinkedIn…

    Early Signs of Rehiring A short update to an earlier post. Big employers from CSX to Alphabet are hiring again after eighteen months of treating hiring as a last resort. The narrow prediction held. Companies cut on the bet that agents would absorb the work, then hired back when the agents did not. The deeper coordination claim stays an open question. Best line, from an MIT economist asked whether firms need more people or fewer: no one has any idea.

    One thread runs through all four. The machine can produce the output. A human still owns the part with no dashboard: the judgment, the handoff, the curb no single car is responsible for. That is my book in one sentence, which is a convenient thing to be able to say now that the fifteen are sitting above.

    Before you go

    Design is where a manuscript stops being a document and starts being a book. I will share the cover here first when it is ready.

    If the book is why you are here, it is Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. Those stories are how I can learn.