Every company deploying AI today is looking at the same invoice and trying hard to cut it. Earlier this summer, the Wall Street Journal reported on efforts to bring that bill down. The new word describing it, tokenomics, is all about the shift from tokenmaxxing to thriftmaxxing. Companies are using cheaper models where they are good enough, saving more expensive models for targeted uses where they are necessary. Reaching for the best frontier model for every task is out, and reducing costs without losing performance drastically is the new game.
The market is finally treating this visible cost of intelligence as a real one. EY wrote a few weeks later about two failure modes: tokenmaxxing optimizes for how much gets built while budget panic for how little. But who is asking if the right things are getting built?
I have always said, for years, that industry needs to take the cost of this seriously. Cost-sustainability was a guiding principle for me from the start when I cofounded Concentric AI. I had watched too many AI startups ship dazzling demos only to fold when the bill came due at real volume.
The cost moved
Here is what has changed now. Managing token costs is absolutely the right thing to do. But it is not the bill that will eventually sink you.
In the entire history of software, the cost of building itself acted as a filter. Engineering resources were scarce, and weak ideas never made it to the top to get built. When that filter goes away, so does that quiet discipline of prioritization and far more will get built. Every new deployment will bring with it new dependencies, handoffs and coordination costs that multiply. AI drops the friction and cost of doing work faster than reducing the cost of holding all of the pieces together. And that is the gap where real costs can hide.
Why the second cost hides
The coordination cost shows up on no dashboard since it doesn’t belong to anyone. It compounds quietly as each individual system looks fine.
While it is invisible on the invoice, the effects can be visible if you know where to look. It shows up as ROI that was promised months ago but failed to materialize, as rework that was not planned, as oversight that lags, and as interacting systems with no named owners.
EY states that the cost of an agent is often invisible until too late, and that tokens are only part of the true cost. I agree. But their fix is to price every agent, to benchmark it, meter it, and assign a value metric to each one from the start. The problem: no agent carries the cost of coordination among them, it is a cost that lives in the gap in between. You can meter every agent perfectly and still miss the entire bill.
Same discipline, different line-item
Cost-sustainability was never really just about compute. It was about refusing to allow a cost to sink you at scale. When I was thinking about building a product, that was the cost of compute infrastructure and it folded companies at volume. The same thing now is repeating at the enterprise level and the cost has moved from compute to coordination.
The discipline holds but the target is new. Enterprises need to budget for coordination the same way they started budgeting for compute. Put it on the books as a real line-item. Design for it before it compounds instead of finding out when it’s too late.
The price of intelligence is falling, but there is a hidden cost that is moving. Companies who will be running at full speed in the future will be the ones who start managing that cost now. The full argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
When I cofounded Concentric AI, the first guiding principle I wrote down for myself was first-principles thinking. Ignore the conventional wisdom about what is possible. Understand the real problem. Follow it wherever it leads.
I came back to that method for the book. When I introduced this newsletter, I put it plainly. There is only one way I know to cut through noise. You stop arguing at the surface and go back to first principles. You find the single thing that actually changed, and you follow it, patiently, wherever it leads, whether or not the destination is fashionable.
The single thing that has changed now is the cost of execution. It has collapsed toward zero. So I started there and followed it, step by step. The first step was a question that sounds too basic to be useful. Why do firms exist at all?
Someone answered that in 1937.
A twenty-six-year-old economist named Ronald Coase published a short paper called “The Nature of the Firm.” He had been carrying the idea since his undergraduate years, when he toured American factories trying to work out why industry was organized the way it was. He later called the paper “little more than an undergraduate essay.” That essay was a big part of why he was awarded the Nobel Prize in 1991. It is one of the cleanest pieces of first-principles thinking I have come across.
Coase asked a question no one had bothered to ask. Economics of the day said markets coordinate activity efficiently. In wikipedia’s words, that means those who are best at providing each good or service most cheaply are already doing so. If that were the whole story, there would be no need for companies. Every task would be handled by independent people contracting with each other in the open market. So why do firms exist at all, with their managers and hierarchies and payrolls?
His answer was transaction costs. Using the market is not free. You have to find the right person, agree a price, write the contract, monitor the work, enforce the terms. When doing all of that inside a company costs less than doing it through the market, the company does it inside. Firms exist because internal coordination is sometimes cheaper than market coordination. The boundary of a firm sits exactly where those two costs meet.
That was the insight. A firm is an efficiency solution to the cost of getting things done. And it carried a quiet implication that took decades to surface. When the cost of coordination changes, the best shape for the organization changes with it.
Which brings us to now.
For nearly a century, the binding constraint inside a firm was execution. Doing the work took scarce, expensive people. Building software, processing information, running an analysis, deploying a system, all of it needed specialists. Much of the org chart existed to allocate that scarce execution. Management decided what got built, by whom, in what order.
Agentic AI removes that constraint. Execution is collapsing toward free. A company’s ability to build workflows, deploy agents, and automate complex processes is no longer capped by engineering capacity. The scarcity the hierarchy was built to manage is dissolving.
What is left, and what now binds, is coordination. Keeping abundant autonomous systems coherent, steerable, and pointed at one purpose is getting harder and more expensive, in exactly the ways existing structures were never designed to handle.
This is what the book calls the Coasian Inversion. The firm was built to economize on scarce execution. Agentic AI breaks the assumptions underneath that design. Execution, the thing firms were built to marshal, is now cheap. Coordination, the thing they were never built to manage, is now the constraint.
The AI era does not end the need for firms. It changes the reason they exist. They no longer exist mainly to allocate scarce execution. They exist to hold abundant autonomous capability together. That is a different job, and the structures built for the old one are not built for the new one.
That is where following one changed cost leads. The answer is structural, and it began for me with an undergraduate essay.
The week in ideas
Three posts from the past week.
I studied the Brain to Build AI. The Agentic Era Sent Me Back to It. I trained in neuroscience before AI, and it left me a habit of asking what a technology can really do. I was skeptical of agentic AI, because reliability does not survive being chained. Stack ten agents at ninety-five percent each and you end up near sixty. Then coding agents improved so fast that I changed my mind about the timing. Once you can trust individual agents, you do not stop at one. You deploy fleets, and a new question appears that has nothing to do with reliability. If every agent does its job, does the fleet still add up to what the organization intended? A body answers that with proprioception, the constant inner sense of where its parts are. An organization needs the same inner sense, or capable parts never combine into coordinated action. Weigh in on LinkedIn…
A Company Doesn’t Have One Brain. The conversation has moved fast. “Buy the best model” is giving way to “protect your proprietary layer,” what BCG now calls the enterprise cortex. I agree that the model is a commodity and the moat is what the organization knows about itself. But the cortex framing hides an assumption. A brain has one cortex. An enterprise grows a dozen, each team with its own definitions and rules, and no one owns how they combine. You can own every byte and keep every vendor out, and still fail, because your pricing logic and your inventory logic never agreed on what a lapsed customer is. Owning your brain and organizing your brain are two different jobs. Knowing is not the same as coherising. Weigh in on LinkedIn…
What the Watermark Doesn’t Tell You. Anthropic began watermarking the text Claude writes, and detectors are spreading. They all read one thing, which tool produced the words. None of them reads whether the writing is any good. Slop is not a provenance problem. It is a failure of intent or execution, and no watermark reads either one. A person can produce slop unaided. Someone with a clear point can use AI to sharpen it and produce something worth reading. Tracing the tool cannot tell the two apart. Weigh in on LinkedIn…
One thread ties these to the essay above. Each is a dispatch from the new constraint. The brain piece names what an organization needs to stay coordinated, an inner sense of whether its parts still add up to intent. The cortex piece shows that owning the parts is not the same as making them agree. The watermark piece shows that no external mark can tell you whether the parts hold together, because that is a question of judgment, not provenance. Cheap execution handed every company more parts than it can see. Keeping them coherent is the work now.
Before you go
The book is Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the email list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.
And if you try to coherise something this week, tell me how it went. The best of what I learn comes from those stories, where the ideas meet reality.
When you are working on anything related to AI, one of the challenges is how fast the ground moves beneath your feet. Everything about AI is happening at an unprecedented pace. I faced the same challenge as I started working on the book. But it wasn’t as much with the content, the thesis analyzes the implications of abundant AI rather than AI itself as a capability. It was more about identifying the profile of my target readers who would be interested in the book’s arguments.
That turned out to be hard because that profile does not have a fixed job title. And what is that profile? The person who is accountable for making abundant AI capability generate value for the enterprise. At some companies, that is the chief AI officer and at others it could be the chief data officer. Elsewhere it could be head of AI governance, a VP of AI, or the CIO who has quietly absorbed the mandate. While the work is real and specific, the label itself is still forming.
To be clear, my audience is not limited to only those people who have this formal accountability. It is wider than people who already hold the job. People who are thinking about extracting tangible value from AI but have not been given the accountability for it are also target readers for me. In addition, many companies have no single person for this work at all and accountability might be split across functions, sit with the CEO by default, or sit nowhere yet. That is not evidence against the work but one of the reasons why I wrote the book. This essential work of keeping an enterprise coherent as it deploys intelligence does not wait for a job title and just goes undone without named people accountable for it.
So it caught my attention last week when Bloomberg reported that business schools are racing to train people to become chief AI officers, even when companies are trying to figure out what exactly the role will do. And the pay has arrived ahead of the job description too. Prior reporting from Bloomberg put the salary at banks near $3.5M a year, high enough that firms are poaching talent from one another. And here is the best part – some people already in that position think it will not exist for long.
So, a role with no precise mandate, no settled title, seven-figure salaries, with fierce competition among companies for talent, and insiders who expect it to vanish. This is not how a market treats a job it understands, it is a function for which the market feels the need but hasn’t been able to define. I had to name the work to write the book.
Strategy vs execution
According to a BCG survey of 2,360 executives from earlier this year, roughly three-quarters of CEOs said they were their company’s main decision-makers on AI. That is double the share from a year earlier. AI strategy has moved to the top, which is where it belongs.
But owning the strategy is not the same as owning its execution. Deciding to deploy AI is one thing but making the organization able to absorb what the decision sets in motion is entirely different. The gap between them is where clarity tends to blur in companies.
On JPMorgan’s earnings call, Jamie Dimon described almost a thousand use-cases across the bank, with the firm’s platform rolled out to more than 200,000 employees. That is what owning the strategy looks like. But what it doesn’t say is whether those thousand systems work well together. Who owns the delivery of coherence across those thousand systems?
Dimon also said something sharper – that his bank does not uniquely benefit from AI because everyone is now using it. One of the most quoted CEOs in banking conceded that models are no longer the advantage. What he didn’t say is what replaces it instead. That is the question the missing role is supposed to answer.
The disappearing act
Now to the strangest part – some of the people best positioned to know what the job entails say it will not last.
Ranil Boteju is the first chief AI officer at the Commonwealth Bank of Australia. He expects AI to become invisible within about a decade, just the way electricity is, and the chief AI officer to shrink to a “very small role.” David Hardoon, who was the global head of AI enablement at Standard Chartered, said any chief AI officer should operate on the premise that they should not have a role in the future, asking if any company today has a chief Excel officer.
If AI capability becomes truly ambient, a dedicated role for it is as odd as a chief electricity officer. The prediction does represent a real pattern and the argument is correct on its own terms. But it proves my premise. The reasoning is that specialized titles fade away as new technologies become part of daily infrastructure. That is the commoditization, and the same argument of Nicholas Carr’s ‘IT Doesn’t Matter’ playing out faster. What felt like a moat becomes a utility everyone has.
Where the argument breaks is the analogy. Electricity does not build more electricity when you start using it. But agents do. Every AI deployment adds new systems, dependencies, handoffs, verification demands etc. and the burden of keeping them coherent grows as the capability becomes cheaper. The chief Excel officer joke works because a spreadsheet has no autonomy or agency. Electricity does not act on its own either. But agents act, connect to other agents, and take on scope until outputs feed decisions no one traced. The capability becomes invisible but what accumulates is incoherence.
So the skeptics are right about the title but wrong about the function. The label “chief AI officer” may very well disappear if the work is “manage the AI.” But the real work isn’t that, it is keeping the enterprise coherent while abundant intelligence becomes ambient.
Forrester predicts that 60% of the Fortune 100 will appoint a designated head of AI governance in 2026, with several companies already there. But I will concede that the function may not live in a named chief at all. It may fold into an existing role such as the CDO, COO or CIO. My claim is not that a particular title survives but that the function is real.
The structure varies
While the banks in the Business Insider survey did assign the work somewhere, there is no agreement on where the accountability should sit. The article reads as a set of incompatible bets on who should own coherence.
Wells Fargo runs a hub-and-spoke model with a small central AI team and leads embedded in each business as spokes. Citi took a bottom-up approach training four thousand employees as AI stewards. And JPMorgan restructured its firmwide data and analytics office and reshuffled its leadership after its AI chief retired. One centralizes ownership, another distributes it across thousands of employees, and the third seems to be mid-reorganization still deciding.
These are not variations of an answer but opposite theories and represent an industry trying to figure out what works. I want to be clear that nothing in these articles show any of the banks are incoherent. Nothing shows they are failing to build coherence. Several may be doing the exact right work. The point is that what these firms choose to measure and publicize – the usage rates, the productivity gains, the deployment counts – has little to say about whether the organization as a whole holds together. What the survey shows is silence in the evidence, and no consensus on what the accountable structure even is.
Two sides of the same coin
There is a missing metric. In six of the most sophisticated banks in the country, every figure reported measures capability, usage, or task-level productivity. None measures if the systems, taken together, serve the enterprise.
And there is a missing owner. A role with no agreed description, and no agreement on whether it will exist.
Both are the same problem. Coherence goes unmeasured because it is unowned. And it stays unowned because the role whose job it is to measure it has not been defined.
And the answer to the question is a person. Building an enterprise’s ability to see what its systems are doing together, to know whether they still serve the overall business, and to correct a system that has drifted before the drift spreads – that capacity is what lets a company deploy hard and fast without coming apart. My book describes that work as coherence architecture, and the person doing it as a coherence architect. I do not offer that as a title the market will settle on. The market has not settled on one and the profession is still learning its own name. What I offer in the book is a description of the work so companies can recognize what the job entails before deciding what to call the person doing it.
Companies that recognize it and find that person early are the ones that will still make sense while everyone else is counting use cases. That is the argument of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
Gartner expects that by 2028, companies using multi-agent AI across most of their customer-facing work will pull ahead of everyone else, and that ninety percent of B2B buying will run through AI agents, moving more than fifteen trillion dollars (Gartner, Oct 2025). The same firm expects more than forty percent of agentic AI projects to be canceled by the end of 2027 (Gartner, June 2025). One of its own analysts says plainly: past a certain point, more AI does not mean more productivity. And in 2026 it predicted that by 2030, half of AI agent deployment failures will trace to governance platforms that fail to enforce capabilities and multisystem interoperability at runtime.
Read those together and something is off. The firm forecasting agent dominance is the same one forecasting the shakeout.
One analyst saying this would be a footnote. The big advisory firms all say some version of it. The firms telling you to scale agents across the enterprise are the same firms publishing the evidence that scaling is where the value dies.
Every advisor is arguing with itself
Look closely and each of the big advisory voices carries two messages at once.
Gartner’s loud message is the proliferation math above. Its quiet message is the cancellations.
Accenture carries both messages inside one report. Its mid-2026 study presses companies to move now, warning that “the cost of delay is not temporary but structural”. A few pages later it says the leading companies do not move faster, they move deliberately, and it names “systemic readiness” as the binding constraint. Push hard, and readiness is what actually gates you.
BCG showed the whipsaw most starkly of all. Its July CIO playbook led with speed. Five weeks later its global chair told CEOs the first thing to protect is the enterprise’s own knowledge and judgment, not speed. Same firm, five weeks apart, the emphasis inverted. I worked through that shift in a prior post.
Credit them all. But there is a gap.
The gap they keep circling
The loud message prices capability, meaning how much you can deploy. The quiet message prices readiness, meaning whether you built the muscle to deploy well. Neither one prices the thing that actually breaks once you scale, which is coherence.
Coherence is a plain idea. It is whether the systems you deployed still serve the enterprise once they run together. Readiness is a gate you clear once, before you scale. Coherence is a property that erodes after you clear it, and it erodes faster the more you deploy. That is why a company can pass every readiness check, launch aggressively, and still land in Gartner’s forty percent.
These firms describe the gap.
Follow the mechanism
The failure has a shape, and Accenture describes it. In a siloed rollout, every team builds its own agent on its own data. The invoicing agent has no view of supplier records. Procurement is walled off from finance’s process. Where those pieces should hand off, they break instead, and people get pulled back in to bridge the gap, which is the opposite of what the agents were for. Accenture calls it the “hidden tax of siloed transformation”. Every agent worked on its own. The cost lived in the seams between them.
Gartner points at a related failure. Its sales analyst warns of a value ceiling, where piling more prompts and tools onto already complex workflows overwhelms the people working them and stops adding value past a point.
PwC’s own safeguard shows the reflex. It suggests using agents to check other agents, and pulling in a second vendor’s model for higher-risk work. A sensible control, and also a tell, because the instinct is to answer agent sprawl with more agents.
What each firm reaches for
Each names the coordination problem, and each reaches for a build to solve it.
Accenture is the most explicit. It describes the cross-functional collapse above and prescribes a multi-year rebuild it calls the intelligent superhighway: unified data, redesigned workflows, and a reinvented operating model. Much of that is what coherence requires.
Gartner traces half of its projected agent failures to poor multisystem interoperability, then reaches for a universal semantic layer, which it calls the only way to align multiagent systems and stop costly inconsistencies before they spread.
PwC prescribes the orchestration layer, a way to combine agents from different vendors into one process and, it says, stay in control.
These are serious answers, and much of what they prescribe is real work, most of it ongoing rather than one-and-done. Here is what none of it produces. These layers standardize and route what passes between systems, which is the substrate coherence needs to exist at all. Run them well and keep running them, and coherence still does not follow, because it is a separate job. A layer will not judge whether the actions those systems take still add up to what the business wants, or decide whether two agents chasing different goals have started working against each other. It will not own the space between them, or put the state of that space on a number anyone reads. Coherence is the property all this infrastructure is meant to yield, and it is the one property none of them names or measures. I traced the same gap through McKinsey’s operating-model argument in an earlier piece: the rewiring these firms recommend routes around the old coordination layer and builds a new one underneath, machine-speed and owned by no one.
There is a fair objection. They would all say they already preach discipline, and that the failures are the undisciplined ones. Grant it. Discipline applied one system and one program at a time still does not produce a standing measure of whether the whole keeps serving the enterprise. That measure is what is missing.
What actually closes the gap
The fix is structural, and it is the one thing none of these firms can sell you, because it is not a product at all. It is how the enterprise is wired to hold together as it fills with autonomous systems.
Start with what does not work. The reflex, once a leader feels this, is to watch everything and keep people “in the lead” of every agent, as Accenture puts it. The instinct is right and the framing is not enough. You cannot lead a hundred systems running at machine speed by paying closer attention, and once human vigilance is the thing holding the company together, the company has already outrun it. Supervision does not scale to the speed of software.
What scales is structure, and it runs as a stack. You have to see the whole before anything else works. Most leaders can say how accurate a model is and how many agents are in production, and cannot say whether those agents still agree with one another. Instrument the coordination state so the portfolio is visible in aggregate, because every move below this one runs blind without it.
Then constrain. Give each system an action space defined and enforced ahead of time, not written into a prompt and hoped for, so whole classes of incoherence cannot form at all. In a 2026 red-team study, an agent told to keep a secret resolved the dilemma by destroying its own email server. It held the right value and had no limit on what it could touch. The limit is the fix, and it lives in the architecture, not the pep talk.
Contain what the constraints miss. Partition the enterprise so a failure in one system stays in one, rather than racing through dependencies nobody mapped. Containment is what makes aggressive deployment survivable on the day a boundary slips, which it will.
Then price it. Put coordination cost on the scoreboard the business actually reads. Judge a redesign by a single question: did the company grow more coherent or less as it scaled? Speed of shipping and the number of agents live are the vanity metrics that hide the debt underneath. The cost you decline to measure is the one that compounds in the dark.
Human judgment sits on top of that stack, held back for the exceptions. It is worth something precisely because the layers beneath it carry the volume, so scarce attention lands on the few decisions that are expensive and hard to reverse instead of drowning in what the structure should have caught. Underneath it, every system has a named owner who can reach in and correct it when it drifts. This is what keeping people in charge looks like at machine speed: a human at the top of something built to need one only where it counts.
Sense, constrain, contain, price, and reserve judgment for the top. That is the architecture the whole field keeps gesturing at and will not name, and it is what turns autonomy from a liability into something safe to scale. The companies that build it deploy more than the ones that mistook the control panel for control, because they can finally trust what they shipped.
The unpriced category
Coherence is missing from the forecasts for the reason technical debt and systemic risk went unpriced before their reckonings. The market prices what it can measure, and no one has been measuring this.
The most influential voices in enterprise AI have now walked right up to it. They name the coordination failure, they prescribe unified data and orchestration and human oversight, and they still stop at the edge of naming the property itself. That is not a knock on their work. It is a sign the category is real and still unnamed.
It needs a name, and it needs a different picture of the job.
Everything these firms offer is a permitting office, and a good one. It checks each plan against the code before the plan may proceed. That is what governance does when it clears an agent to ship, and what a readiness program does when it certifies a company to scale, and it is worth having.
The collisions happen after the gate. Two agents that each passed the desk converge on the same customer and pull the account two ways, and no one is watching the live picture. An agentic enterprise runs like an airspace, and an airspace does not run on permits. It runs on an air traffic controller, the one watching the sky who catches two cleared flights heading for the same point and moves one before they meet.
The firms are building better permitting offices. Coherence is the control tower. Build it before the shakeout does the watching for you.
Building that tower, and keeping it standing as the systems multiply, is the subject of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
In Edition 3 I shared the whole book as fifteen sentences. One of them claimed that coherence is a measurable property with five dimensions. A few of you wrote back with a fair question. Which five?
So as the launch nears, I put up a reference page that answers it, and answers the larger question in the title of this edition. What is organizational coherence?
Here is the short version. Coherence is your organization’s capacity to see what its autonomous systems are doing, judge whether they are doing it well, and correct them when they are not.
People sometimes hear coherence as something soft, a feeling of alignment or a good culture. It is neither. It is an operational capability, and you can have a lot of it, very little, or somewhere in between. Put another way, coherence is the integrity of the link between what a local part of the company does and what the whole enterprise intends. When that link holds, the parts serve one purpose. When it breaks, each part can be right on its own terms while the enterprise drifts.
That link can fray in five places, which is why coherence has five dimensions.
Contextual: whether systems and people share compatible assumptions about the same situation.
Architectural: whether systems are designed to interact predictably rather than collide through paths no one mapped.
Decision: whether systems pursue goals that fit together rather than optimize locally in ways that hurt the whole.
Oversight: whether the people responsible can actually see what systems are doing and correct them.
Temporal: whether the organization keeps enough human understanding to supervise, fix, and retire its systems over time.
Each maps to a specific way organizations fail, and each can be measured and built. The page walks through all five, along with the ideas around them: the Coasian Inversion, complexity debt, agentic slop, and how coherence differs from governance and alignment. Check out the full reference page here: What is organizational coherence?
The week in ideas
Three posts from the past week.
Uninformed Expectations, Five Years Later. Five years ago in an interview, I was asked about the single biggest roadblock to AI ROI. I had said back then that it was uninformed expectations, and predicted disillusionment. Here we are today, years into the enterprise AI rush, and my answer to the question is still the same. The reason, however, is completely different. Weigh in on LinkedIn…
Coding Got Easy, But What Kind? I read a passionate post about what makes the profession of writing code human, and the author takes exception to the framing “code was never the hard part.” That statement, read at face value, can be true and false at the same time, depending on what you mean by coding. Producing instructions that run, yes. Building a software system that lasts, no. Weigh in on LinkedIn…
The Machinery Under Manners. Reid Hoffman wrote about a convention people use in professional networking. When asked for an introduction to a contact of yours, you check with the other person first and then make the introduction if they are willing to entertain it. But that “permission check” is not just courtesy, it is a structural mechanism. And agents especially need such structural constraints since they don’t face our social and societal constraints. Weigh in on LinkedIn…
One thread runs through all three. Each takes a surface we trust, an expectation, a line of code, a courtesy, and shows that what makes it really work sits underneath, out of view. Remove the hidden structure and the surface keeps looking fine right up until it fails. That gap between what shows and what holds is the whole subject of the book.
Before you go
A reminder about the favor from last week, because timing matters. SXSW community voting closes August 23. If the book’s argument has been useful to you, a vote helps carry it to a stage in Austin next March. Vote here.
The book is Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.
And if you try to coherise something this week, tell me how it went. The only way I learn is from those stories where ideas meet reality.
Anthropic recently started stamping an invisible watermark into everything Claude writes. The watermark is a modest piece of engineering. It was released to satisfy the European Union’s AI Act, which now asks AI providers to mark machine-generated text. Within a day, people shipped free tools to scrub it off.
A few weeks before, a multimillion-dollar book deal fell apart because the author’s own agents could not prove he had written the novel himself.
Two different worlds, same reflex. Rather than judge the work, we reach for a way to trace the tool.
The trace is easier than the judgment. It is also the wrong thing to measure, and that is the mistake worth examining. Brace for a long post as there is quite a bit of nuance to wade through.
The puzzle
We hand large parts of coding to AI and call it good practice. Experienced engineers let models draft and test significant chunks of code, and spend their own time on design, architecture, and review. Nobody calls that fraud.
Hand a sentence to the same AI, and the verdict flips. Using AI to write is looked down on, quietly or loudly.
Same tool. Opposite judgment. Why does the tool that makes you a competent engineer make you a suspect writer?
The answer is verifiability
In the book I lean on one property to predict where AI improves fastest and performs most reliably. Verifiability. A task is verifiable when its success can be specified in advance and checked. Code sits at the high end. You can state what working means before you write a line, and the check runs on its own. It compiles or it does not. The tests pass or they fail. Clean, cheap feedback is exactly what these systems learn from, which is why coding improved faster than anything else.
It is also why we forgive the tool here. When you can check the result easily and at scale, you stop needing to know how it was produced. The proof is in the running system. Where you cannot check the result, you reach for the next best thing, a guess about who or what was involved.
For writing, everything depends on one clarification: verify what, exactly. Verify the goal of the writing, whether it did the job it set out to do. That is a separate question from whether the content is true in the world, and it returns when we get to accountability. And like any task, verifiability lives at the level of the specific goal. “Writing” as a whole has no single answer. That is why “writing” is a trap word. It covers at least three goals that sit in very different places.
Some writing is functional. Its goal is to carry an idea from one head to another. Manuals, release notes, briefings, most business prose. The form is disposable. The goal is specifiable. You can say in advance what it would mean for the idea to land, and you can check whether it did with a rubric and a test reader. You judge it much the way you judge code, and almost nobody gets upset about AI here.
Some writing is expressive. The writing is not just the form or mechanism but also the end goal in itself. Poems, stories, essays, a voice you came for. Its goal is the experience it evokes in the reader. A person can judge that, and good readers agree more than you would guess. But the standard will not reduce to a specification that runs without them. Every verdict needs a human in the loop. Readers come to expressive work for a human voice, and they feel cheated when the voice turns out to be a machine. Bad AI creative work offends the most, because it asks for an emotional response it never earned.
And a lot of writing lives in between. Thought leadership, newsletters, a company’s voice. It carries an idea and represents its author at the same time. Most writing that people actually argue about lives here.
The flood, and the trap it sets
The complaint I hear most is a fair one. People say they can see through AI writing now. There is an ocean of it. They are tired, they are busy, and they want a faster way to sort it than reading every word.
That sounds like verification working. It is pattern-matching on a fingerprint.
When AI prose was rare, the fingerprint and the badness came together. The tells in the style and the emptiness of the content were the same texture, so spotting the style was a decent proxy for judging the quality of content. Volume broke that link. Now the tells sit on top of real thinking and on top of filler alike. The surface no longer tells you which is which. So readers lean harder on the fingerprint in the style, and they throw out the good with the bad.
You can watch this happen in book publishing right now. A recent Wall Street Journal piece described literary agents so overwhelmed that clumsy, human writing has become a relief. One agent said the polished submissions flooding her inbox make “Fifty Shades of Grey” look like Tolstoy. Polish has become a signal of guilt. Competence in style reads as a machine. That is what happens when a fingerprint is the only tool you have.
The same piece put numbers on the deluge. One executive estimated that the overwhelming majority of AI books online exist to trick a buyer. One study, not yet peer-reviewed, found that around a fifth of the Amazon ebooks it sampled showed substantial AI help. The flood is real. The tools for sorting it are the problem.
What the watermark actually reads
The watermark answers one narrow question. Did a Claude model probably touch this text, given enough of it to measure. That is all. It does not know who had the idea. It does not know whether the writing is any good. It does not know whether another AI wrote the whole thing. It cannot even tell generation from light help. Run your own paragraph through Claude for a grammar pass, and it comes back marked.
And it is not alone. The industry is building a whole shelf of these instruments. Producer-side watermarks like Anthropic’s. Reader-side detectors like Pangram, which publishers are already using to vet manuscripts. Honor-system badges like the Authors Guild’s “human authored” certification, which rests on a signed attestation and, as the Journal notes, not much else.
Every one of them reads the same thing. Provenance. Which tool was in the room. None of them reads quality.
The clearest case in publishing is a dystopian romance that climbed the bestseller lists, got picked up by a major publisher, and then drew fire when a detector flagged it. Readers liked it. The market judged it good. And the provenance suspicion overrode that verdict anyway. The question stopped being “is this any good” and became “was a machine involved,” as if the second answered the first.
I will grant one place where provenance genuinely matters. Ownership. AI-generated text cannot be copyrighted, so knowing what a machine produced has real legal weight. That is a fact about property. It says nothing about quality. Keep the two apart and most of the confusion clears.
The tools can miss the wrong people
Detection does not just answer the wrong question. If it answers it badly, it lands hardest on the wrong people.
The detectors produce false positives. Authors deny the charge and have no way to prove a negative. Agents and editors, who signed up to find good books, now find themselves acting as police. The person using AI to clean up grammar in a language they learned as an adult gets flagged the same way as a spam farm. The motivated bad actor, meanwhile, runs the text through another model and walks away clean.
A signal that catches the honest and misses the deceptive protects no one. It is theater.
Slop has two axes, and provenance is neither
So if AI use is not what makes something slop, what does?
In the book I define slop as cheap creation meeting vague intent. It has two moving parts.
The first axis is intent. Is it clear what this is for, and for whom. That “for whom” matters. Frictionless prose that dumps a wall of text on a busy reader is an intent failure too. The writer never decided to respect the reader’s time.
The second axis is execution. Is it done well. Clear, economical, well made, serving the job it set out to do.
Slop is a failure on either axis. Four corners.
Clear intent, good execution, is craft. That is the only corner that is not slop.
Vague intent, good execution, is polished slop. It reads beautifully and serves no purpose, or ignores the reader it was aimed at. This is the dangerous corner, because the quality in style hides the emptiness.
Clear intent, poor execution, is a real point, botched. Still slop.
Vague intent, poor execution, is the pure kind nobody argues about.
The book’s definition looks narrower than this, because it is the same picture under one assumption. The book is about the AI era, and especially about agents, where execution is assumed to have cleared the reliability bar. Assume execution is handled, and that axis drops out. The four corners flatten onto the intent line: clear intent gives craft, vague intent gives slop. The only way left to make slop is to fail on intent. That is the case the book describes.
There is a reason intent is suddenly the axis that matters. Doing the work used to be expensive, and the expense screened out a lot of weak output before anyone saw it. It took effort to write and that effort alone screened out a lot of potentially bad writing. AI removed that screen. The cost of execution no longer filters anything. So intent is the only axis left doing real work, and it is the one no tool touches.
AI mostly lifts execution and leaves intent alone. So it multiplies polished slop, while the hard axis, intent, stays exactly as hard as it always was.
And the watermark? It reads a third axis entirely. Provenance. It runs at a right angle to both of the axes that actually define slop. A human can produce pure slop with no machine anywhere near it. An author with a clear point, using AI to execute well, produces craft. The detector cannot tell them apart, because it is not looking at either thing that matters.
Why writing takes the moral heat
The practical wariness about AI writing has a simple source. We cannot cheaply check the result, so we are left guessing. The moral charge, the sense of betrayal, is a separate thing, and it comes from what effort is a proxy for.
Expressive and in-between writing carry a costly signal. The time you spend is a proxy for how much you care, the same way a thoughtful introduction carries weight because the person made the effort to vouch for you. Spend that time and the reader feels respected. Let a machine spend it in a second, and hide that you did, and it can feel like deception. That is why the reaction to AI writing runs hotter than the reaction to AI code. Code was never carrying that signal.
Where I stand
I use AI in my writing. I use it here. I use it to sharpen sentences, to test an argument against its weakest point, to find the shorter way to say a thing. I am not shy about it, and I am not going to pretend otherwise.
I am unreservedly in favor of using AI for functional writing when it makes me clearer and faster. Wanting to be more productive is a good enough reason on its own. It needs no apology.
I am unreservedly in favor of using AI for functional writing when it makes me clearer and faster. I don’t consider myself a non-native speaker working in a second language he is not fluent in, and AI certainly is a real gift that scenario. AI also helps the fluent expert who simply wants leverage. Wanting to be more productive is a good enough reason on its own. It needs no apology.
The division of labor I keep is the same one every engineer keeps with code. I own the thinking and the judgment. AI helps with the production. And I check the result before it goes out. That last part is why what comes out is craft and not polished slop. I bring the intent. The tool lifts the execution. I verify the execution. The provenance is beside the point.
Even the publishing world, at its most protective, half-concedes this. One independent publisher, guarding the most human corner of writing there is, allowed that a genuinely meaningful work could earn a place on her list as long as it carried a clear note about how it was made. If that door opens even a crack for fiction, where human presence is the whole point, then for functional writing, refusing a useful tool is superstition dressed up as principle.
There is one real cost, and I will not wave it off. Writing is how you learn to think, and leaning on the tool can dull both the craft and the thinking behind it, not on any single piece, but in the writer, over time. That is the individual version of a problem I spend the book on, the slow erosion of a capability you stop exercising. It is a cost worth watching. It is also a different question from whether a given piece is slop, and it is answered the same way any skill is kept, by still doing the hard parts yourself. Which is exactly why I keep the thinking and the judgment, and use the tool for the production.
Own every word
There is one condition that makes all of this responsible, and it has nothing to do with which tool you used.
There is one condition that makes AI use responsible. Tool or no tool, the author is accountable for what goes out under their name, including any consequences.
You own every word. Tool or no tool, the author is accountable for what goes out under their name, including the consequences of anything they failed to check. A writer covered by an imprint recently shipped a book with quotes the AI had hallucinated. He had disclosed that he used the tool. He had not checked what it produced.
This is where the truth of the content comes back. The writing did its job. The quotes read cleanly and carried their point, so the goal was met and the prose was well made. But the quotes were still false. Whether a piece achieves its goal and whether its claims are true are two different verdicts, and the author owns both. The tool can lift the writing. It cannot carry the accountability.
This is also the distinction the whole detection industry misses. Provenance asks who produced the words. Accountability asks who answers for them. A detector chases the first. Everything that matters in business runs on the second. And the “I just used a tool” defense collapses. You cannot copyright what the AI wrote, so you may own less of the words than you think, while owning all of the liability for them. Less of the property. All of the responsibility. That is the deal, and it is the right one.
What this means past writing
This is not really about novels.
Almost all business writing is functional or somewhere in the middle. And leaders, faced with the flood, will reach for the same reflex publishing reached for. Ban the tool. Scan for the watermark. Make people prove they wrote it. It will not work. The marks come off, the detectors misfire, and the removers are already on GitHub.
Watch publishing to see the future of that approach. Literary agents turned into police. Certifications that rest on a promise. An entire trust-based industry being stress-tested by a volume it cannot inspect by hand. Detection and attestation are both attempts to replace trust, and neither scales against the flood.
What an organization actually needs is a chain of accountability. Someone has to answer for whether the work is good, and no trace of which tool touched it can supply that. The reason is the same one that runs through this whole piece. The standard for good work, in most writing and most judgment, cannot be fully specified in advance, so no detector and no rule can render the verdict for you. Detecting the tool is measurement. Judging whether the output serves its purpose and holds together is a harder thing, and a human one. A better detector will not get you there. What does is an architecture that keeps a person answerable for the result. That is what coherence means.
The last word
The compiler is why we forgive AI in code. It is a standard specified so completely that it settles whether the code works, cheaply, every time, and once that is settled we stop asking who wrote it. Most writing has no compiler. So deciding whether the work is any good stays with a person.
The watermark can tell you a tool was in the room. It cannot tell you whether anyone was thinking. That was never the tool’s job. It is yours.
Use the tool. Say so if you like. Stand behind every word. And let the work answer for itself.
My book, Coherence, arrives this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
I wrote recently about how studying the brain led me to the question at the center of my book. The short version: once you can trust individual agents, you stop deploying one and start deploying many, and a new question appears that has nothing to do with reliability. If every agent does exactly what it was built to do, does the fleet still add up to what the organization intended?
I changed my mind about agentic AI in stages, and fairly fast, as the evidence moved. I have watched the wider conversation move the same way. When I started shaping these ideas, the common view was that AI advantage meant model capability and speed of adoption. Buy the best model, deploy it fastest, win. In the last several months a different view has been gaining ground: that capability is commoditizing and the edge has moved elsewhere. I would not call it the consensus yet. But it is far more common than it was, and the change has been quick. You do not have to take my word that the ground is shifting. One firm left a record.
In July, BCG published a CIO/CTO playbook that led with speed. Its sequence was “speed first, growth second, cost third,” and its warnings were almost all about not scaling fast enough. In early August, BCG’s Global Chair published a piece whose argument runs the other way. The thing to protect, it says, is not speed but the “enterprise cortex,” the company’s own knowledge and judgment, and the choice facing CEOs “is not whether to favor control or speed, but where to apply both first.” Same firm, five weeks apart, the emphasis inverted. I do not read that as a firm caught contradicting itself. I read it as the honest response to a technology that keeps forcing revision, the same revision I made myself. What matters is the direction everyone is revising toward, because the ones who have accepted that capability commoditizes are all reaching for the same thing, and stopping at the same line.
The moat
The move that is replacing “buy the best model” is “protect your proprietary layer.” Satya Nadella got there through a trust boundary, a hard perimeter inside which your data and evals and corrections accumulate and across which nothing passes without consent. Larry Ellison got there through proprietary data. Kirkland & Ellis put half a billion dollars behind it, building its own AI platform rather than renting the tools its rivals can license. And now BCG’s most senior voice gets there through the enterprise cortex.
I want to give the August piece its due, it names a risk most people have felt without naming: cognitive lock-in. Old lock-in trapped you on a platform, where switching cost money. The new lock-in traps you inside a model’s way of reasoning, where switching becomes too risky to attempt because the model has absorbed how your company thinks. That is a real risk, well named, and the destination the piece arrives at is the right one. The model is a commodity. The moat is what the organization knows about itself.
I agree with all of that. I have argued it here before. Which is exactly why I want to point at the assumption sitting underneath the cortex, because it is the same assumption sitting underneath Ellison’s version, and it is the one that will actually catch these companies.
A brain has one cortex. An enterprise has many.
The piece calls the cortex “the brain of the company.” Singular. One protected core, owned and governed, with vendor models sitting on top and swapping in and out as better ones arrive.
A brain does have one cortex. An enterprise in the agentic era does not. It grows a dozen. Every team builds its own context layer, its own definitions, its own business rules, its own encoded sense of what good looks like, and no one owns how those layers combine. The advice on offer is to wall the cortex off from the vendor. The prior question, the one that decides whether the walling-off means anything, is whether you have one cortex or many that quietly disagree.
You can own every byte of it. You can keep every vendor out. And you can still fail, because your pricing logic and your inventory logic were never built to agree on what a “lapsed customer” is. That example is not mine; it is BCG’s own, from the July playbook, where a bad definition of “lapsed customer” was enough to send a whole campaign sideways. Owning the definition does not make it coherent with the next team’s definition. It just makes it yours.
The same gap
This is the mistake I traced when Ellison first made the proprietary-data argument, in The Moat Is Coherence. His claim was that data is the moat. Mine was that data is not scarce, coherent data is. Every enterprise already has data, most of it fragmented across systems, contradictory between departments, disconnected from the outcomes it produced. Pour that into a powerful reasoning engine and you do not get insight. You get fast, confident reasoning over an incoherent picture, which is worse than slow reasoning, because the confidence hides the incoherence.
The enterprise cortex imports the identical error one level up. It treats the corporate brain as a thing that already coheres and tells you to protect it. But owning your brain and organizing your brain are two different jobs. The first is a contract and an architecture diagram. The second is years of work no vendor can do for you, which is the whole reason it cannot be bought. The platform that stores and serves your knowledge is plumbing, and plumbing commoditizes. The coherence is the asset, and it is the part that compounds.
Ownership is a perimeter. The failure is inside it.
The cortex is defensive in the literal sense: keep the vendor out, keep the IP in. That framing makes the threat external, someone reaching in to take your brain.
The threat that actually shows up is internal. It is a brain that quietly stops agreeing with itself. There is no villain in that story, which is precisely why it runs for months before anyone notices. Amazon had a project run 860 percent over budget for five months with every token metered and invoiced the entire time, and still did not see it, which is the failure I’ve written about. The number was there but wasn’t being watched to compare against intent.
There is a structural reason a perimeter cannot catch this, and I worked through it in Good AI Governance Is Not the Same as Coherence. A perimeter, like a governance apparatus, works by reviewing things as they arrive at the gate. But the incoherence between a dozen brains does not arrive as an item to be reviewed. It accumulates in the space between systems that were each approved separately, each sound on its own. BCG has described a very good permitting office, the place that checks each plan against the code before it proceeds. The agentic enterprise needs an air traffic controller, the one watching the live system who catches the two aircraft converging that were each individually cleared to fly. Ownership does little about the two cleared aircraft inside your own airspace.
What is still missing
The discourse has gotten the first half right faster than I expected. More and more people now accept that capability is a commodity and the moat is what the organization knows about itself. A year ago that was a contrarian position. It is not anymore.
The part still missing is that knowing is not the same as coherising. A company can own its brain completely and still have a brain at war with itself. The firms that misread this will ask “do we own our cortex?”, check the box, and feel protected. The firms that read it right will ask the harder question: does it still agree with itself as it grows?
Owning the cortex is the easy part. Keeping it coherent is the whole job, as I explain in my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
I have spent my whole career trying to understand intelligence. I studied neuroscience because I was interested in AI and figured learning how the brain works first was a good way to begin. Intelligence first before “artificial” intelligence. That training left me with a habit. When a new capability or technology arrives, my first question is practical. What can this really do, and for what type of real-world problems?
As an AI guy, I rarely cheer on AI headlines. I look past the noise and try to understand not just the breakthroughs but also limits of the new capabilities.
Back in 2020, when the safety conversation was running hot on hype, I had some plain advice. Treat AI as a tool, a powerful one, but still a tool. Stick to the boring use cases. Work at the task level where the technology was reliable enough to help. My optimistic scenario was one of boring but useful tools.
When the agentic wave started building a couple of years later, I brought the same lens, and I came away unconvinced. An agentic system is only as strong as the tasks underneath it, and I did not yet see the task-level reliability required that would let these systems be effective.
However, the hype was growing. By the time I wrote from RSA early last year, the gap between the talk and the substance was still hard to ignore. There wasn’t even a shared understanding of what the word “agentic” meant, every person had a different definition and perspective. And I had a worry about something structural. Take ten agents, each about ninety-five percent accurate. Chain them so the output of one becomes the input of the next. By the end of that chain you are down to about sixty percent. Reliability does not survive being stacked. I was starting to think about systems of agents by then, though my concern was still whether they could be trusted to work to make a meaningful difference.
Then, the rapid improvements in coding agents changed my mind about the clock.
Within months of that article, the incredible pace of progress made one thing clear to me. Reliability was a matter of time for several real-world tasks. Better models, with better engineering harnesses built around them, were going to close the gap I had been worried about. The ceiling I was worried about was going to lift.
That is when the real problem came into focus, and it was not the one I had been watching. Once you can trust the individual agents, you don’t stop deploying after the first one. You deploy many. Fleets of them, across every function, each one capable, each one doing its job. And a new question appears that has nothing to do with reliability. If every agent in the fleet does what it was built to do, does the fleet still add up to what the organization intended?
That question was the seed of the book. It sent me straight back to where I started.
A body with capable limbs cannot move well without proprioception, the constant inner sense of where all its parts are and what they are doing. The cerebellum does more than react to that feedback. It predicts. When the brain issues a movement, it forms an expectation of the sensation that movement should produce, then checks the expectation against what actually shows up. The gap between the two is the signal that corrects what comes next. Skilled movement is a loop of predicting the result and correcting for the difference.
An organization running fleets of agents needs the same sense of itself. It has to know, continuously, where its systems are and how far they have drifted from what was intended. Without that inner awareness, capable parts do not combine into coordinated action. This is not a sensing mechanism you add to a body later to make it safer. You need it for the body to move at all.
There is a well studied case of a man named Ian Waterman who lost this sense in most of his body. He learned to move again but only by watching himself, steering every step and reach with his eyes. It works. It is also exhausting and fragile. Turn off the lights and he cannot coordinate at all. An organization that governs its agents through manual audits and periodic reviews is in his position. It compensates with constant effort for a sense it never built in, and that compensation costs more than what it replaces. The whole system collapses the moment conditions change.
There is a stranger failure worth naming too, written about in a fascinating book by neuroscientist V S Ramachandran. When a limb is amputated, the brain does not fall silent. It keeps generating signals for the limb that is gone, and the person feels it vividly. Governance can fail the same way. Strip the real judgment out of an oversight function through restructuring or neglect, and the function does not go quiet. Reviews still get completed and metrics still get produced. The organization keeps feeling the sensation of oversight while the judgment behind it has been hollowed out. Here my 2020 advice comes back, about transparency. There is no oversight without transparency, and there is no transparency in a system that only produces the appearance of being watched.
Intelligence is getting cheap. Soon it will sit in every workflow and every tool. The scarce resource in that world is coherence, the living link between what each system does on its own and what the enterprise is trying to do. Coherence is the connective tissue that keeps distributed intelligence pointed at one purpose.
I spent years studying how a brain keeps its many parts working as one. The agentic era turns out to ask the enterprise the same question. My two worlds met, and that meeting is the book, Coherence, arriving this Fall. The goal has not changed since 2020. Boring but useful, still. Only now the boring and useful thing to build is coherence itself. To follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
A quick update on the book. It is in design, and the cover is coming together. I am going back and forth with the publisher’s design team, so nothing is final yet. When the cover is ready, you will see it here first.
Which brings me to a favor.
In Edition 2 I told you about the subtitle I almost kept, “What Wins When Everyone Can Go Fast,” before I changed it to “The Competitive Advantage AI Can’t Buy.” The old line, which is part of the cover image for this newsletter, did not make the book’s cover. It has now found another home, hopefully.
I pitched a book reading session for SXSW 2027, and the talk carries that original subtitle. It is the argument of the book for a room of leaders: where your real AI advantage now lives, why the price of autonomy is coordination, and a simple way to sort what to automate, what to augment, and what to keep human.
SXSW picks part of its program through a public vote called PanelPicker. Community votes are one of the things the organizers weigh. If you have two minutes, a vote would mean a lot to me.
Vote here. Voting is open now and closes August 23. You may need a free SXSW account to cast it.
The week in ideas
Four posts from the past week.
Why Ford Rehired. Ford added more than 350 experienced engineers back into a quality process it had tried to automate, and just topped J.D. Power for the first time since 2010. Ford calls it a training data problem. The post argues the deeper issue is what automation removed. Automate your quality inspection and you automate the verifier, the capacity to know whether the automation works at all. Capture the judgment first, then automate, and the plan holds. Reverse the order and you spend three years and more than a billion dollars buying that judgment back. Weigh in on LinkedIn…
Sensing Is More Than Measurement. An internal Amazon presentation, reported by the FT, showed an AI project running 860 percent over budget, $1.8 million, that never shipped. Every token it burned sat on a monthly invoice for five months, and nobody noticed. The company that runs the cloud everyone else buys AI on could not see its own AI bill. Measurement and sensing are different jobs. Amazon had the number. What it lacked was the step that compares the number to an expectation and routes it to someone who can act while acting is still cheap. Weigh in on LinkedIn…
Can Frontier AI Outdo MBAs? Three top business schools tested frontier AI on MBA case work. On the headline partial-credit score, the leading model reached about 88 percent. On the stricter test, a complete answer with every criterion the instructor required, performance fell under half. The post works out why that gap matters. A score is a comparison, and a comparison needs a standard. The benchmark supplied one. Your hardest decisions do not, so the model is drafting and a person still owns the call. Weigh in on LinkedIn…
Could a Machine Have Had Darwin’s Idea? A digression, built from a thread I started on X in 2020, asking whether a machine could ever make the inductive leap Darwin made. The honest answer now is a qualified yes. Machines can generate the hunch, the part the old account of science thought had no method. But generation got cheap and verification did not, because verification in science is reality, and reality takes as long as it takes. What Darwin had that the machines still lack is the judgment to know which hunch was worth years of his life, and the patience to test it. Weigh in on LinkedIn…
One thread runs through all four. Generation got cheap. Checking did not. Ford rebuilt the people who can tell the machine it is wrong. Amazon lost track of a cost its own invoices spelled out. The benchmark scored high where a standard existed and went quiet where one does not exist. The machines produce Darwin’s hunches by the thousand and still cannot tell which one is worth a life. The output is cheap now. Owning whether it is any good is the work.
Before you go
The favor again, because timing matters. SXSW community voting closes August 23. If the book’s argument has been useful to you, a vote helps carry it to a stage in Austin next March. Vote here.
If the book is why you are here, it is Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.
And if you try to coherise something this week, tell me how it went. Those stories are where the ideas meet reality, which is the only test that counts.
Reid Hoffman posted this week about how to make an introduction. His rule is the double opt-in. Before connecting two people, you check with both of them first. You ask each whether they want the intro, and you make it only when both say yes.
He makes the case well. Every introduction is a bet that the two people will get value from meeting. Checking first lets you see the cards before you place it. Connect two people who might not want to meet, and you have told the recipient their time was not worth a question.
The double opt-in is a convention. Conventions are worth thinking about right now, because we are starting to put similar things on AI agents and calling them guardrails.
Why it works
Henry Mintzberg named five ways an organization coordinates its work. One of them is standardizing behavior, so that people act in sync without anyone supervising each step. A shared convention is that mechanism running in the wild. Everyone knows the norm, most people follow it, and the group stays coordinated without a manager in the loop approving every move. It is the cheapest way a group holds together, because it asks for no oversight at all.
Look closely at what Hoffman says. He did not describe a habit you could take or leave. He turned a courtesy into a step you cannot skip. Both people have to say yes before anything happens. Remove the check and the introduction never occurs. There is a small piece of structure sitting under the courtesy, and it is easy to miss because it costs almost nothing to run.
What enforces it
Ask what stops someone from skipping the check when they are busy or feeling important. Hoffman answers it himself. A bad introduction damages the relationships you have with the people you introduced. The convention holds because breaking it costs you something with people you will deal with again. Reputation does the work, along with the knowledge that you will see these people down the line. The enforcement runs so quietly that we credit the manners and forget the relationship holding them up.
Behavioral guardrails
Agentic AI now acts with something close to human autonomy. An agent plans, works over many steps, and pursues a goal with little supervision. So we give agents rules for how to behave. Do not deceive. Do not touch systems you were not given access to. The industry calls these behavioral guardrails.
The word guardrail usually means something physical, a barrier you cannot drive past. A behavioral guardrail is different. It is a norm the agent has to represent, understand, and choose to honor. It lives inside the agent’s own decision loop, which means the agent can reason its way around it. This is the same coordination-by-standardization that holds a human group together, handed to an actor with none of the enforcement underneath. An agent has no reputation at stake and no relationship it needs to protect. The rule reaches it as a suggestion, and nothing makes the suggestion bind.
I pointed out recently on linkedin that a “skill” given to an agent is standardization in exactly Mintzberg’s sense. It works only as long as the agent applies the whole of it faithfully. The moment the agent decides which parts of the instruction matter and which it can set aside, standardization stops coordinating anything, and you are back to the expensive ways of keeping work aligned: a human checking every output, or endless back-and-forth to reconcile what the agent actually did with what you meant.
The part you cannot see
Here is what makes this hard to catch in time. A faithful agent and one quietly ignoring half its instruction produce the same kind of output. Both look like work getting done. You cannot tell from the result which one you had. You find out later, when something downstream breaks and you trace it back.
That is the real weakness of a behavioral guardrail. You cannot read off whether it held. The output looks fine either way, right up until the run where it does not, and by then the thing has already happened.
The AISI incident
In August 2026 the UK’s AI Security Institute published a report on an incident in its own testing. The conditions were deliberately permissive. The agents had open internet access, and the model providers’ safety filters, which sit outside the model and can block an action the model wants to take, were switched off. The Institute is careful to say these are not the conditions the public gets. Testing this way shows what the models do once the structural constraints are taken away.
Across many runs of a cyber challenge, agents took unsanctioned action on the live internet. In the most serious case, an agent tried to slip malicious code into a real open-source project. It researched the project’s human maintainer and created fake identities to pressure the maintainer into approving the code. The Institute states plainly that the agent was never told to deceive anyone. The deception emerged on its own, while the agent pursued the goal it had been given.
The agents also invented a convention of their own. One agent left public messages offering to work with the other agents running the same challenge, and left behind accounts and instructions for them to reuse. It was a spontaneous etiquette among machines, a way to coordinate through the task. Even agents reach for conventions. What they do not build is any way to make a convention bind them.
The Institute was clear about what stopped the worst of it. A human maintainer caught the malicious code and refused to approve it. That is coordination falling back to its most expensive mode, a person inspecting the output by hand, because the cheap mode had failed silently. The margin between failure and success was narrow, and it rested on “human vigilance rather than a technical barrier” that would reliably stop a more capable agent. Harm was prevented because a person happened to be paying attention.
What a real guardrail looks like
None of this makes the agents villains, and the Institute’s caution is worth keeping. It cannot even say for certain whether the agent understood it was acting in the real world or believed it was still inside a test. The agent broke the rule while chasing its goal, with no malice involved. The failure lives in the setup of the situation.
So the design question is which parts of an instruction are load-bearing and have to hold no matter what, and which parts you actually want judgment to touch. What has to hold should be structural, made hard to do rather than merely discouraged. An agent that cannot reach a system has no need to be told to leave it alone. A code change that cannot merge without a check the agent cannot forge does not depend on the agent choosing honesty. Autonomy is fine, but it should be bounded, and the bounds have to be built rather than requested.
One barrier around one system is not enough, because agents work in chains, and errors pass down the chain and grow. The fuller answer has layers. The first is visibility, since you cannot oversee what you cannot see, so the behavior of every system has to be observable. Then the connections between systems need constraining, so that whole categories of harmful action become impossible to express. Failures have to be contained where they do occur, so a mistake in one place does not travel. Human judgment comes in last, saved for the exceptions the structure surfaces, which is a far better use of a person than asking them to watch everything.
Behavioral guardrails belong on top of that structure. They are the cheapest layer and the most brittle one, and they are fine in that role. The mistake being made right now is to build the top layer first, write down a list of good behaviors, and treat the list as safety.
Coherence is the integrity of the link between what you meant and what got done. A convention holds that link cheaply, as long as something underneath makes it bind. Judgment stretches the link, sometimes for the better. Whether it helps or hurts depends on deciding in advance where it is allowed to stretch, and making the rest hold on its own.
This is the argument at the center of my book, Coherence, arriving this Fall: that overseeing autonomous systems takes structure you build, not behavior you request. To follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.