Category: The Argument

Why execution stopped being scarce, and why coherence became the constraint. Maps to Part One of the book “Coherence.”

  • What the Watermark Doesn’t Tell You

    Anthropic recently started stamping an invisible watermark into everything Claude writes. The watermark is a modest piece of engineering. It was released to satisfy the European Union’s AI Act, which now asks AI providers to mark machine-generated text. Within a day, people shipped free tools to scrub it off.

    A few weeks before, a multimillion-dollar book deal fell apart because the author’s own agents could not prove he had written the novel himself.

    Two different worlds, same reflex. Rather than judge the work, we reach for a way to trace the tool.

    The trace is easier than the judgment. It is also the wrong thing to measure, and that is the mistake worth examining. Brace for a long post as there is quite a bit of nuance to wade through.

    The puzzle

    We hand large parts of coding to AI and call it good practice. Experienced engineers let models draft and test significant chunks of code, and spend their own time on design, architecture, and review. Nobody calls that fraud.

    Hand a sentence to the same AI, and the verdict flips. Using AI to write is looked down on, quietly or loudly.

    Same tool. Opposite judgment. Why does the tool that makes you a competent engineer make you a suspect writer?

    The answer is verifiability

    In the book I lean on one property to predict where AI improves fastest and performs most reliably. Verifiability. A task is verifiable when its success can be specified in advance and checked. Code sits at the high end. You can state what working means before you write a line, and the check runs on its own. It compiles or it does not. The tests pass or they fail. Clean, cheap feedback is exactly what these systems learn from, which is why coding improved faster than anything else.

    It is also why we forgive the tool here. When you can check the result easily and at scale, you stop needing to know how it was produced. The proof is in the running system. Where you cannot check the result, you reach for the next best thing, a guess about who or what was involved.

    For writing, everything depends on one clarification: verify what, exactly. Verify the goal of the writing, whether it did the job it set out to do. That is a separate question from whether the content is true in the world, and it returns when we get to accountability. And like any task, verifiability lives at the level of the specific goal. “Writing” as a whole has no single answer. That is why “writing” is a trap word. It covers at least three goals that sit in very different places.

    Some writing is functional. Its goal is to carry an idea from one head to another. Manuals, release notes, briefings, most business prose. The form is disposable. The goal is specifiable. You can say in advance what it would mean for the idea to land, and you can check whether it did with a rubric and a test reader. You judge it much the way you judge code, and almost nobody gets upset about AI here.

    Some writing is expressive. The writing is not just the form or mechanism but also the end goal in itself. Poems, stories, essays, a voice you came for. Its goal is the experience it evokes in the reader. A person can judge that, and good readers agree more than you would guess. But the standard will not reduce to a specification that runs without them. Every verdict needs a human in the loop. Readers come to expressive work for a human voice, and they feel cheated when the voice turns out to be a machine. Bad AI creative work offends the most, because it asks for an emotional response it never earned.

    And a lot of writing lives in between. Thought leadership, newsletters, a company’s voice. It carries an idea and represents its author at the same time. Most writing that people actually argue about lives here.

    The flood, and the trap it sets

    The complaint I hear most is a fair one. People say they can see through AI writing now. There is an ocean of it. They are tired, they are busy, and they want a faster way to sort it than reading every word.

    That sounds like verification working. It is pattern-matching on a fingerprint.

    When AI prose was rare, the fingerprint and the badness came together. The tells in the style and the emptiness of the content were the same texture, so spotting the style was a decent proxy for judging the quality of content. Volume broke that link. Now the tells sit on top of real thinking and on top of filler alike. The surface no longer tells you which is which. So readers lean harder on the fingerprint in the style, and they throw out the good with the bad.

    You can watch this happen in book publishing right now. A recent Wall Street Journal piece described literary agents so overwhelmed that clumsy, human writing has become a relief. One agent said the polished submissions flooding her inbox make “Fifty Shades of Grey” look like Tolstoy. Polish has become a signal of guilt. Competence in style reads as a machine. That is what happens when a fingerprint is the only tool you have.

    The same piece put numbers on the deluge. One executive estimated that the overwhelming majority of AI books online exist to trick a buyer. One study, not yet peer-reviewed, found that around a fifth of the Amazon ebooks it sampled showed substantial AI help. The flood is real. The tools for sorting it are the problem.

    What the watermark actually reads

    The watermark answers one narrow question. Did a Claude model probably touch this text, given enough of it to measure. That is all. It does not know who had the idea. It does not know whether the writing is any good. It does not know whether another AI wrote the whole thing. It cannot even tell generation from light help. Run your own paragraph through Claude for a grammar pass, and it comes back marked.

    And it is not alone. The industry is building a whole shelf of these instruments. Producer-side watermarks like Anthropic’s. Reader-side detectors like Pangram, which publishers are already using to vet manuscripts. Honor-system badges like the Authors Guild’s “human authored” certification, which rests on a signed attestation and, as the Journal notes, not much else.

    Every one of them reads the same thing. Provenance. Which tool was in the room. None of them reads quality.

    The clearest case in publishing is a dystopian romance that climbed the bestseller lists, got picked up by a major publisher, and then drew fire when a detector flagged it. Readers liked it. The market judged it good. And the provenance suspicion overrode that verdict anyway. The question stopped being “is this any good” and became “was a machine involved,” as if the second answered the first.

    I will grant one place where provenance genuinely matters. Ownership. AI-generated text cannot be copyrighted, so knowing what a machine produced has real legal weight. That is a fact about property. It says nothing about quality. Keep the two apart and most of the confusion clears.

    The tools can miss the wrong people

    Detection does not just answer the wrong question. If it answers it badly, it lands hardest on the wrong people.

    The detectors produce false positives. Authors deny the charge and have no way to prove a negative. Agents and editors, who signed up to find good books, now find themselves acting as police. The person using AI to clean up grammar in a language they learned as an adult gets flagged the same way as a spam farm. The motivated bad actor, meanwhile, runs the text through another model and walks away clean.

    A signal that catches the honest and misses the deceptive protects no one. It is theater.

    Slop has two axes, and provenance is neither

    So if AI use is not what makes something slop, what does?

    In the book I define slop as cheap creation meeting vague intent. It has two moving parts.

    The first axis is intent. Is it clear what this is for, and for whom. That “for whom” matters. Frictionless prose that dumps a wall of text on a busy reader is an intent failure too. The writer never decided to respect the reader’s time.

    The second axis is execution. Is it done well. Clear, economical, well made, serving the job it set out to do.

    Slop is a failure on either axis. Four corners.

    Clear intent, good execution, is craft. That is the only corner that is not slop.

    Vague intent, good execution, is polished slop. It reads beautifully and serves no purpose, or ignores the reader it was aimed at. This is the dangerous corner, because the quality in style hides the emptiness.

    Clear intent, poor execution, is a real point, botched. Still slop.

    Vague intent, poor execution, is the pure kind nobody argues about.

    The book’s definition looks narrower than this, because it is the same picture under one assumption. The book is about the AI era, and especially about agents, where execution is assumed to have cleared the reliability bar. Assume execution is handled, and that axis drops out. The four corners flatten onto the intent line: clear intent gives craft, vague intent gives slop. The only way left to make slop is to fail on intent. That is the case the book describes.

    There is a reason intent is suddenly the axis that matters. Doing the work used to be expensive, and the expense screened out a lot of weak output before anyone saw it. It took effort to write and that effort alone screened out a lot of potentially bad writing. AI removed that screen. The cost of execution no longer filters anything. So intent is the only axis left doing real work, and it is the one no tool touches.

    AI mostly lifts execution and leaves intent alone. So it multiplies polished slop, while the hard axis, intent, stays exactly as hard as it always was.

    And the watermark? It reads a third axis entirely. Provenance. It runs at a right angle to both of the axes that actually define slop. A human can produce pure slop with no machine anywhere near it. An author with a clear point, using AI to execute well, produces craft. The detector cannot tell them apart, because it is not looking at either thing that matters.

    Why writing takes the moral heat

    The practical wariness about AI writing has a simple source. We cannot cheaply check the result, so we are left guessing. The moral charge, the sense of betrayal, is a separate thing, and it comes from what effort is a proxy for.

    Expressive and in-between writing carry a costly signal. The time you spend is a proxy for how much you care, the same way a thoughtful introduction carries weight because the person made the effort to vouch for you. Spend that time and the reader feels respected. Let a machine spend it in a second, and hide that you did, and it can feel like deception. That is why the reaction to AI writing runs hotter than the reaction to AI code. Code was never carrying that signal.

    Where I stand

    I use AI in my writing. I use it here. I use it to sharpen sentences, to test an argument against its weakest point, to find the shorter way to say a thing. I am not shy about it, and I am not going to pretend otherwise.

    I am unreservedly in favor of using AI for functional writing when it makes me clearer and faster. Wanting to be more productive is a good enough reason on its own. It needs no apology.

    I am unreservedly in favor of using AI for functional writing when it makes me clearer and faster. I don’t consider myself a non-native speaker working in a second language he is not fluent in, and AI certainly is a real gift that scenario. AI also helps the fluent expert who simply wants leverage. Wanting to be more productive is a good enough reason on its own. It needs no apology.

    The division of labor I keep is the same one every engineer keeps with code. I own the thinking and the judgment. AI helps with the production. And I check the result before it goes out. That last part is why what comes out is craft and not polished slop. I bring the intent. The tool lifts the execution. I verify the execution. The provenance is beside the point.

    Even the publishing world, at its most protective, half-concedes this. One independent publisher, guarding the most human corner of writing there is, allowed that a genuinely meaningful work could earn a place on her list as long as it carried a clear note about how it was made. If that door opens even a crack for fiction, where human presence is the whole point, then for functional writing, refusing a useful tool is superstition dressed up as principle.

    There is one real cost, and I will not wave it off. Writing is how you learn to think, and leaning on the tool can dull both the craft and the thinking behind it, not on any single piece, but in the writer, over time. That is the individual version of a problem I spend the book on, the slow erosion of a capability you stop exercising. It is a cost worth watching. It is also a different question from whether a given piece is slop, and it is answered the same way any skill is kept, by still doing the hard parts yourself. Which is exactly why I keep the thinking and the judgment, and use the tool for the production.

    Own every word

    There is one condition that makes all of this responsible, and it has nothing to do with which tool you used.

    There is one condition that makes AI use responsible. Tool or no tool, the author is accountable for what goes out under their name, including any consequences.

    You own every word. Tool or no tool, the author is accountable for what goes out under their name, including the consequences of anything they failed to check. A writer covered by an imprint recently shipped a book with quotes the AI had hallucinated. He had disclosed that he used the tool. He had not checked what it produced.

    This is where the truth of the content comes back. The writing did its job. The quotes read cleanly and carried their point, so the goal was met and the prose was well made. But the quotes were still false. Whether a piece achieves its goal and whether its claims are true are two different verdicts, and the author owns both. The tool can lift the writing. It cannot carry the accountability.

    This is also the distinction the whole detection industry misses. Provenance asks who produced the words. Accountability asks who answers for them. A detector chases the first. Everything that matters in business runs on the second. And the “I just used a tool” defense collapses. You cannot copyright what the AI wrote, so you may own less of the words than you think, while owning all of the liability for them. Less of the property. All of the responsibility. That is the deal, and it is the right one.

    What this means past writing

    This is not really about novels.

    Almost all business writing is functional or somewhere in the middle. And leaders, faced with the flood, will reach for the same reflex publishing reached for. Ban the tool. Scan for the watermark. Make people prove they wrote it. It will not work. The marks come off, the detectors misfire, and the removers are already on GitHub.

    Watch publishing to see the future of that approach. Literary agents turned into police. Certifications that rest on a promise. An entire trust-based industry being stress-tested by a volume it cannot inspect by hand. Detection and attestation are both attempts to replace trust, and neither scales against the flood.

    What an organization actually needs is a chain of accountability. Someone has to answer for whether the work is good, and no trace of which tool touched it can supply that. The reason is the same one that runs through this whole piece. The standard for good work, in most writing and most judgment, cannot be fully specified in advance, so no detector and no rule can render the verdict for you. Detecting the tool is measurement. Judging whether the output serves its purpose and holds together is a harder thing, and a human one. A better detector will not get you there. What does is an architecture that keeps a person answerable for the result. That is what coherence means.

    The last word

    The compiler is why we forgive AI in code. It is a standard specified so completely that it settles whether the code works, cheaply, every time, and once that is settled we stop asking who wrote it. Most writing has no compiler. So deciding whether the work is any good stays with a person.

    The watermark can tell you a tool was in the room. It cannot tell you whether anyone was thinking. That was never the tool’s job. It is yours.

    Use the tool. Say so if you like. Stand behind every word. And let the work answer for itself.

    My book, Coherence, arrives this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • A Company Doesn’t Have One Brain

    I wrote recently about how studying the brain led me to the question at the center of my book. The short version: once you can trust individual agents, you stop deploying one and start deploying many, and a new question appears that has nothing to do with reliability. If every agent does exactly what it was built to do, does the fleet still add up to what the organization intended?

    I changed my mind about agentic AI in stages, and fairly fast, as the evidence moved. I have watched the wider conversation move the same way. When I started shaping these ideas, the common view was that AI advantage meant model capability and speed of adoption. Buy the best model, deploy it fastest, win. In the last several months a different view has been gaining ground: that capability is commoditizing and the edge has moved elsewhere. I would not call it the consensus yet. But it is far more common than it was, and the change has been quick. You do not have to take my word that the ground is shifting. One firm left a record.

    In July, BCG published a CIO/CTO playbook that led with speed. Its sequence was “speed first, growth second, cost third,” and its warnings were almost all about not scaling fast enough. In early August, BCG’s Global Chair published a piece whose argument runs the other way. The thing to protect, it says, is not speed but the “enterprise cortex,” the company’s own knowledge and judgment, and the choice facing CEOs “is not whether to favor control or speed, but where to apply both first.” Same firm, five weeks apart, the emphasis inverted. I do not read that as a firm caught contradicting itself. I read it as the honest response to a technology that keeps forcing revision, the same revision I made myself. What matters is the direction everyone is revising toward, because the ones who have accepted that capability commoditizes are all reaching for the same thing, and stopping at the same line.

    The moat

    The move that is replacing “buy the best model” is “protect your proprietary layer.” Satya Nadella got there through a trust boundary, a hard perimeter inside which your data and evals and corrections accumulate and across which nothing passes without consent. Larry Ellison got there through proprietary data. Kirkland & Ellis put half a billion dollars behind it, building its own AI platform rather than renting the tools its rivals can license. And now BCG’s most senior voice gets there through the enterprise cortex.

    I want to give the August piece its due, it names a risk most people have felt without naming: cognitive lock-in. Old lock-in trapped you on a platform, where switching cost money. The new lock-in traps you inside a model’s way of reasoning, where switching becomes too risky to attempt because the model has absorbed how your company thinks. That is a real risk, well named, and the destination the piece arrives at is the right one. The model is a commodity. The moat is what the organization knows about itself.

    I agree with all of that. I have argued it here before. Which is exactly why I want to point at the assumption sitting underneath the cortex, because it is the same assumption sitting underneath Ellison’s version, and it is the one that will actually catch these companies.

    A brain has one cortex. An enterprise has many.

    The piece calls the cortex “the brain of the company.” Singular. One protected core, owned and governed, with vendor models sitting on top and swapping in and out as better ones arrive.

    A brain does have one cortex. An enterprise in the agentic era does not. It grows a dozen. Every team builds its own context layer, its own definitions, its own business rules, its own encoded sense of what good looks like, and no one owns how those layers combine. The advice on offer is to wall the cortex off from the vendor. The prior question, the one that decides whether the walling-off means anything, is whether you have one cortex or many that quietly disagree.

    You can own every byte of it. You can keep every vendor out. And you can still fail, because your pricing logic and your inventory logic were never built to agree on what a “lapsed customer” is. That example is not mine; it is BCG’s own, from the July playbook, where a bad definition of “lapsed customer” was enough to send a whole campaign sideways. Owning the definition does not make it coherent with the next team’s definition. It just makes it yours.

    The same gap

    This is the mistake I traced when Ellison first made the proprietary-data argument, in The Moat Is Coherence. His claim was that data is the moat. Mine was that data is not scarce, coherent data is. Every enterprise already has data, most of it fragmented across systems, contradictory between departments, disconnected from the outcomes it produced. Pour that into a powerful reasoning engine and you do not get insight. You get fast, confident reasoning over an incoherent picture, which is worse than slow reasoning, because the confidence hides the incoherence.

    The enterprise cortex imports the identical error one level up. It treats the corporate brain as a thing that already coheres and tells you to protect it. But owning your brain and organizing your brain are two different jobs. The first is a contract and an architecture diagram. The second is years of work no vendor can do for you, which is the whole reason it cannot be bought. The platform that stores and serves your knowledge is plumbing, and plumbing commoditizes. The coherence is the asset, and it is the part that compounds.

    Ownership is a perimeter. The failure is inside it.

    The cortex is defensive in the literal sense: keep the vendor out, keep the IP in. That framing makes the threat external, someone reaching in to take your brain.

    The threat that actually shows up is internal. It is a brain that quietly stops agreeing with itself. There is no villain in that story, which is precisely why it runs for months before anyone notices. Amazon had a project run 860 percent over budget for five months with every token metered and invoiced the entire time, and still did not see it, which is the failure I’ve written about. The number was there but wasn’t being watched to compare against intent.

    There is a structural reason a perimeter cannot catch this, and I worked through it in Good AI Governance Is Not the Same as Coherence. A perimeter, like a governance apparatus, works by reviewing things as they arrive at the gate. But the incoherence between a dozen brains does not arrive as an item to be reviewed. It accumulates in the space between systems that were each approved separately, each sound on its own. BCG has described a very good permitting office, the place that checks each plan against the code before it proceeds. The agentic enterprise needs an air traffic controller, the one watching the live system who catches the two aircraft converging that were each individually cleared to fly. Ownership does little about the two cleared aircraft inside your own airspace.

    What is still missing

    The discourse has gotten the first half right faster than I expected. More and more people now accept that capability is a commodity and the moat is what the organization knows about itself. A year ago that was a contrarian position. It is not anymore.

    The part still missing is that knowing is not the same as coherising. A company can own its brain completely and still have a brain at war with itself. The firms that misread this will ask “do we own our cortex?”, check the box, and feel protected. The firms that read it right will ask the harder question: does it still agree with itself as it grows?

    Owning the cortex is the easy part. Keeping it coherent is the whole job, as I explain in my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • I studied the Brain to Build AI. The Agentic Era Sent Me Back to It.

    I have spent my whole career trying to understand intelligence. I studied neuroscience because I was interested in AI and figured learning how the brain works first was a good way to begin. Intelligence first before “artificial” intelligence. That training left me with a habit. When a new capability or technology arrives, my first question is practical. What can this really do, and for what type of real-world problems?

    As an AI guy, I rarely cheer on AI headlines. I look past the noise and try to understand not just the breakthroughs but also limits of the new capabilities.

    Back in 2020, when the safety conversation was running hot on hype, I had some plain advice. Treat AI as a tool, a powerful one, but still a tool. Stick to the boring use cases. Work at the task level where the technology was reliable enough to help. My optimistic scenario was one of boring but useful tools.

    When the agentic wave started building a couple of years later, I brought the same lens, and I came away unconvinced. An agentic system is only as strong as the tasks underneath it, and I did not yet see the task-level reliability required that would let these systems be effective.

    However, the hype was growing. By the time I wrote from RSA early last year, the gap between the talk and the substance was still hard to ignore. There wasn’t even a shared understanding of what the word “agentic” meant, every person had a different definition and perspective. And I had a worry about something structural. Take ten agents, each about ninety-five percent accurate. Chain them so the output of one becomes the input of the next. By the end of that chain you are down to about sixty percent. Reliability does not survive being stacked. I was starting to think about systems of agents by then, though my concern was still whether they could be trusted to work to make a meaningful difference.

    Then, the rapid improvements in coding agents changed my mind about the clock.

    Within months of that article, the incredible pace of progress made one thing clear to me. Reliability was a matter of time for several real-world tasks. Better models, with better engineering harnesses built around them, were going to close the gap I had been worried about. The ceiling I was worried about was going to lift.

    That is when the real problem came into focus, and it was not the one I had been watching. Once you can trust the individual agents, you don’t stop deploying after the first one. You deploy many. Fleets of them, across every function, each one capable, each one doing its job. And a new question appears that has nothing to do with reliability. If every agent in the fleet does what it was built to do, does the fleet still add up to what the organization intended?

    That question was the seed of the book. It sent me straight back to where I started.

    A body with capable limbs cannot move well without proprioception, the constant inner sense of where all its parts are and what they are doing. The cerebellum does more than react to that feedback. It predicts. When the brain issues a movement, it forms an expectation of the sensation that movement should produce, then checks the expectation against what actually shows up. The gap between the two is the signal that corrects what comes next. Skilled movement is a loop of predicting the result and correcting for the difference.

    An organization running fleets of agents needs the same sense of itself. It has to know, continuously, where its systems are and how far they have drifted from what was intended. Without that inner awareness, capable parts do not combine into coordinated action. This is not a sensing mechanism you add to a body later to make it safer. You need it for the body to move at all.

    There is a well studied case of a man named Ian Waterman who lost this sense in most of his body. He learned to move again but only by watching himself, steering every step and reach with his eyes. It works. It is also exhausting and fragile. Turn off the lights and he cannot coordinate at all. An organization that governs its agents through manual audits and periodic reviews is in his position. It compensates with constant effort for a sense it never built in, and that compensation costs more than what it replaces. The whole system collapses the moment conditions change.

    There is a stranger failure worth naming too, written about in a fascinating book by neuroscientist V S Ramachandran. When a limb is amputated, the brain does not fall silent. It keeps generating signals for the limb that is gone, and the person feels it vividly. Governance can fail the same way. Strip the real judgment out of an oversight function through restructuring or neglect, and the function does not go quiet. Reviews still get completed and metrics still get produced. The organization keeps feeling the sensation of oversight while the judgment behind it has been hollowed out. Here my 2020 advice comes back, about transparency. There is no oversight without transparency, and there is no transparency in a system that only produces the appearance of being watched.

    Intelligence is getting cheap. Soon it will sit in every workflow and every tool. The scarce resource in that world is coherence, the living link between what each system does on its own and what the enterprise is trying to do. Coherence is the connective tissue that keeps distributed intelligence pointed at one purpose.

    I spent years studying how a brain keeps its many parts working as one. The agentic era turns out to ask the enterprise the same question. My two worlds met, and that meeting is the book, Coherence, arriving this Fall. The goal has not changed since 2020. Boring but useful, still. Only now the boring and useful thing to build is coherence itself. To follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Coding Got Easy, But What Kind?

    A programmer named Senko Rašić published an angry post this week. He is angry at a slogan going around: “code was never the hard part.” He calls it an insult to programmers. The post hit Hacker News and Lobsters and drew more than five hundred comments.

    Reading the comments, I noticed one thing over and over. They argue about what the word “coding” means. Some people say it is the easy part and always was. Others say it is the whole job and always was. They are not disagreeing about difficulty. They are using one word for two different things: producing code, and building something that holds.

    I have a stake in this. I write that execution has gotten cheap and coherence is the hard thing now. Skim that fast and you might file me under the same slogan.

    The hard part was always building well

    Producing code that runs is one thing. Building something that holds is another: the design that survives real data, the structure that stays coherent as the system grows, the choices that still make sense a year later when the author is gone. That second thing was always the hard part, and it was always the job.

    Before AI, doing it well was expensive, because it took a skilled person and their time. Doing it badly was possible, but it was slow and the result was bad. Nothing about producing software was cheap.

    Cheap is what AI added. A model writes a working function or a working page in the time it takes to describe it. Someone with no training can now get running code out of a sentence. That is new, and I will not soften it. Execution got cheap.

    Cheap is not the same as bad

    AI writes good code in places, and the commenters who said so are right. The quality is uneven, and the unevenness has a shape. The models are strong where they had the most to imitate, the patterns written out in public a million times: CRUD, forms, glue code, the endpoint that reads the database and returns JSON. They are weaker where the examples run out, on scientific computing, embedded work, anything performance-critical. One commenter put it well: AI kills it on problems with a thousand forum posts, and you do not point it at the ten-billion-dollar machine headed to Mars.

    There is a reason coding is where AI advanced fastest. Code is checkable. It runs or it does not, the tests pass or they do not, and my book argues that AI improves fastest exactly where the work can be checked. The checkable parts of building will keep getting cheaper and better. What stays hard is the judgment no test can catch.

    That was never the job

    I wrote recently about a delete button from a data engineering thread. An engineer built a button to remove a record. It took the record off the screen. It left behind the three related records the original had created when it was made. The screen looked right. The data underneath was orphaned.

    Producing that button is the cheap part. A model does it now. Knowing it had to clean up three records you cannot see is the judgment. That was the hard part, and that was the job.

    The comparison people keep reaching for

    Watching the threads, I noticed people reaching for the same analogy – writing. Anyone can put words into sentences that flow and reach a point. A model does it fluently. That was never what made someone a writer. What makes a writer is the choice of words and their order. The same point can land or die on those choices. One commenter compared it to a novel, where clean sentences were never the hard part.

    AI did not invent bad prose but made fluent-looking prose free and endless. “Code was never the hard part” runs the same move as saying words were never the hard part of writing. About spelling, it is a shrug. About writing, it is an insult. The slogan gets both out of the same words by letting you hear the first while it means the second.

    The swap, named

    The slogan is true about producing code and false about building well. It earns its credibility on the first and spends it on the second, where the conclusion is that coders are now optional. One commenter worried the line would harden into a truism for business leaders. That is the reader I have in mind.

    If someone read me as saying coders no longer matter, they would be making the same swap. The thing I say got cheap is producing code. The thing I call hard is building something that holds. Those did not both get cheaper. Anyone worried about systems built fast and badly is saying that building them well is still hard. You cannot write about that problem and also believe the work is trivial.

    The right version of the slogan is my argument

    Some people say the line and do mean something reasonable. On Lobsters, a commenter called lcamtuf said Senko was reading it uncharitably, and that the real meaning is that producing lines of code was never the bottleneck. That is true. If coding was fifteen percent of an engineer’s week, automating all of it buys back about a sixth of the week. You aimed the speedup at the fastest part of the job. Another commenter reached for Amdahl’s Law to make the same point with a number.

    Then lcamtuf added the part that matters most to me. Some friction, he said, was good, because it stopped people from building software that was unnecessary or unmaintainable. That friction was a filter. It did not stop people from producing code. It stopped building-without-judgment from becoming load-bearing, because getting anything shipped had to pass through people whose time was scarce. That filter is the thing AI removed. The difficulty of building well did not leave. It stopped being enforced.

    The same thing, at two heights

    Building well has a name once you zoom out. Making each part fit a whole you cannot see all of at once is coherence.

    The delete button is a coherence failure inside one function. The record fit the screen and broke the data underneath. Agentic slop is the same failure across a company. Each workflow works on its own, and the enterprise stops making sense. Senko is defending coherence in one program. I am after it across an organization.

    The slogan is backwards. The hard part of code was never the typing. It was always the judgment, and here is the part that lasts. Where a design was already worked out a thousand times in public, the model has something to copy, and it copies well. Where the design is new to your system, there is nothing to copy and no test to guide it. That judgment stays human, because it is both uncheckable and unwritten.

    Producing code got cheap. That judgment never did. It used to get paid up front, in salaries and review and the time of people who knew what they were doing, where you could see the cost. Now it shows up late, in the corners, with no owner.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Uninformed Expectations, Five Years Later

    In January 2021 an interviewer asked me for the biggest roadblock to AI adoption. My answer came down to one thing: expectations. I called them uninformed. Organizations believed AI was a drop-in component that would improve whatever process it touched. That belief, mixed with a fear of missing out, pushed companies to rush. I wrote that they were reaching for “a new shiny hammer looking for nails,” and that disillusionment would follow.

    I still think that was right. The disillusionment arrived on schedule. But I was also wrong.

    Two versions of one mistake

    The 2021 belief was easy to state. Buy the model and the value follows. AI was a part you slotted into an existing workflow. Reality punished that belief fast, because building anything real was hard. You needed engineering time, data science help, and a business case strong enough to justify the spend. Projects that underestimated the work usually stalled before they shipped. The return died at the front of the pipeline, where the building happened.

    The belief has since turned inside out. Building is easy now. A product manager can stand up an agent in an afternoon with no engineering queue in sight. So the new expectation is that easy building means easy value. If a workflow takes a day instead of a quarter, the returns should take care of themselves.

    They don’t. And the reason traces back to the same root as before.

    Capability was never the outcome

    Both beliefs make the same move. They treat capability as if it were the outcome. Capability is what your AI can do. The outcome is a separate thing: what your organization still has once the AI has done it. The space between those two is where the return leaks away.

    In 2021 that space was easy to see, because scarcity kept it visible. To build anything, a team had to win scarce engineering time. Winning it meant convincing people outside the team. A budget owner. An architect. A security reviewer. Nobody designed that as oversight. It was just the price of a scarce resource. It still worked like a filter. Weak ideas died in the queue, and only the ones someone could defend reached production. Scarcity was doing quiet work that never showed up on an org chart.

    That filter is gone. When building costs almost nothing, nothing stops a weak idea from becoming a running system. Fifty teams can each ship their own agent, every one of them green on its own dashboard, and no single person owns the question of what they add up to. The return still leaks. It leaks at the far end of the pipeline now, in the cost of coordinating systems nobody mapped.

    The newest version of the old belief

    What is the belief being sold right now? The frontier model vendors have told the market that the gap to ROI is expertise. You have the models. What you lack is people who know how to wire them into your environment. So the labs send forward deployed engineers. They embed at your site, build the integrations, tune the configurations, debug the odd behavior, and leave a working system behind.

    The role exists because my diagnosis is right. Building the model was never the hard part. Deploying it inside a messy enterprise is. FDEs are a real answer to that, and a good one. They are also an answer to the wrong problem, and their structure guarantees it.

    Start with the incentive. An FDE works for the vendor. Success for them means adoption and a satisfied customer. They have no reason, and usually no mandate, to tell you that a deployment conflicts with a system three departments away that they cannot see, or that it will cost you more in complexity than it returns in efficiency, or that the right call is to not build it. The most valuable act of coherence is sometimes the word no. You cannot buy that from the party paid to say yes.

    Then there is what they can see. An FDE embedded in one business unit knows that unit. They have no view of the other agents running across the company, or the coordination surfaces their new system quietly creates. The failures I worry about do not come from one bad deployment. They come from the sum of many reasonable ones. No FDE is positioned to see the sum.

    And there is what they leave behind. When the engagement ends, the deepest understanding of why the system was built that way, what it assumes, and how to change it safely often leaves with them. You inherit a running system and a dependency, not the knowledge to oversee it over time.

    So the FDE belief is the 2021 belief again, dressed for 2026. In 2021 it was “buy the model and value follows.” Now it is “add the deployment experts and value follows.” Both stop at capability. Deployment velocity is still capability. It is the thing every competitor can rent from the same labs, on the same terms, in the same quarter. It gets the system live. It does not decide whether the system should have gone live at all, and it does not hold the enterprise together once fifty of them are running. FDEs solve half the problem. The half they cannot touch is the one that decides your return.

    Why the old advice still works

    Back then I offered a few tips. Go slow. Pick narrow, well-defined use cases. Get a quick win on the board. Kill a project when the evidence says to, and keep sunk cost from making that call for you. I stand behind every word of it. What surprises me is why it still holds.

    In 2021 that discipline was prudence. Ignore it and scarcity would punish you, so the advice helped you get through the queue with something that worked. Today the same discipline is close to the only filter left in the building. “Go slow” used to save you from a stalled project. Now it is most of what stands between you and a sprawl of systems you cannot see. “Kill the project” used to fight sunk cost. Now it fights the agent that keeps running because nobody confirmed it should stop.

    The advice is unchanged. What used to enforce it for free has disappeared, so the work now falls to you. Scarcity handled a crude version of this by accident. You have to handle the real version on purpose.

    The name I was missing

    When I wrote that post I was reaching for something I could not name. I knew rushing was dangerous. I knew AI was a means to an end. I had no word for the property that separates a company that gets value from one that gets debt.

    The word is coherence. It is the capacity to see what your autonomous systems are doing, judge whether they are doing it well, and correct them when they are not. In 2021, scarcity supplied a crude version of it by accident. In 2026, you build it on purpose or you go without. Once execution gets cheap, coherence becomes the thing that decides whether all that abundance turns into advantage or into cleanup.

    The disillusionment I flagged five years ago is still on its way to a lot of organizations. It reaches them from the opposite direction now. Back then it came from systems that were too hard to build. Now it comes from systems that are too easy to build. The belief underneath is the one I named in 2021. Stop mistaking what AI can do for what your organization will keep, and build the coherence that turns the first into the second.

    The full argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Could a Machine Have Had Darwin’s Idea?

    In November 2020 I started a thread on X. I have barely used the platform these past couple of years, and I only came back to this thread recently, by accident, when some of my own old posts surfaced in front of me. Reading them in one sitting was strange, because a thread I had added to piecemeal over five years turned out to have been about one question the whole time.

    It opened with Darwin. Someone I followed had posted, amazed, about what evolution manages to build, and I wrote that what amazed me was something else: the intellectual leap required to infer the theory in the first place, from nothing but empirical observation and years of focused study. I said I was not sure I fully appreciated Darwin’s intellect and dedication. Then, in the next post, I said the thing the whole thread has been chasing ever since. That kind of leap, I wrote, is the kind of intelligence AI should be aiming for, not recognizing cat faces or copying tasks people already do well.

    That was the bar I set in 2020, and it was a high one. Darwin spent decades buried in finches and barnacles and pigeon breeders and fossil beds, a mountain of unconnected observation, and then made a leap: one idea, natural selection, that reorganized all of it at once. That inductive move, from a heap of messy particulars to the principle that explains them, struck me then as the most formidable act an intellect can perform, and the part of real science I was least sure a machine could touch. So I kept a list, adding to it whenever I came across a case of AI looking like it was helping advance science rather than just crunch it, to watch whether anything ever cleared the bar.

    Reading the thread back now, it traces an arc I did not plan, and it worried at the right problem from the start. Within days, in November 2020, I posted a New York Times piece on Max Tegmark’s group, whose neural network had recovered a hundred physics equations from raw data, and I pulled out the catch that the physicists themselves named. Tegmark was candid that the machine could retrieve the formulas but not yet the deep principles beneath them, the quantum uncertainty or the relativity that would explain why the formula holds. And Jesse Thaler, the MIT physicist directing the new AI-and-physics institute, put his finger on why. AI wins at games because a game has a well-defined notion of success. “If we could define what success means for physical laws,” he said, “that would be an incredible breakthrough.” Proposing the theory, in other words, was the hard part, and it was hard precisely because you could not say in advance what would count as getting it right.

    The entries that moved me most in those early days were not about AI at all. They were about humans doing the thing I wanted to see a machine do. Around the same time I was reading about Tibor Gánti, the Hungarian biologist who deduced from first principles what the simplest possible living thing must be: a metabolism, a way to store information, and a membrane, three systems that have to be coupled or the organism dies. I wrote at the time that this was the kind of inductive thinking we should be teaching children. Later I followed a related idea, assembly theory, which proposes a single elegant handle on complexity: count the minimum number of steps needed to build a molecule, and use that number to test for the presence of life on other worlds. A theory reduced to one measurable quantity, aimed at one of the hardest questions there is. It may not survive contact with the evidence; I later noted that a study found some minerals scoring above the threshold the theory sets for life. But right or wrong, it was the move I admired, the leap from observation to a principle sharp enough to be tested.

    Then the machine examples accumulated. In December 2021, a Nature paper where neural networks guided the intuition of mathematicians toward new conjectures in knot theory, the machine surfacing the pattern and the human still making the leap. In March 2022, a system that rediscovered Newton’s law of gravitation from the motion of the planets and wrote it back out as a symbolic equation. Then more symbolic regression, pulling laws from data. Then, in February 2025, the entries that made me sit up: Google’s AI co-scientist, generating and ranking novel research hypotheses on its own, and Evo-2, a model that does not just read genomes but writes them. And I drew a line I did not fully understand at the time. I noted that the most exciting work was different from the systems that merely “generate plausible hypotheses from an extremely large space of possibilities.”

    Five years of watching, and that distinction turned out to be the whole story.

    The step that was supposed to have no method

    For a century, the standard account of science drew a hard line between two acts.

    One is coming up with the idea. The other is checking whether the idea is true. Karl Popper named these the context of discovery and the context of justification, and he was blunt about which one belonged to philosophy. Testing a hypothesis has a logic. Having one does not. In his words, the act of conceiving a theory “neither calls for logical analysis nor is susceptible of it.” The initial leap from a pile of observations to this might be why was, he thought, a matter for psychology, not method. A hunch. The unteachable part.

    That is where the whole romance of science lived, and Darwin’s leap is its patron saint. Kekulé dreaming the benzene ring as a snake biting its tail. Fleming noticing the one culture plate that had gone wrong in an interesting way. Darwin holding twenty years of specimens in his head until they resolved into a single idea. We told these stories because the leap seemed to come from nowhere, and coming from nowhere was the point. You could train someone to run an experiment. You could not train the hunch.

    The systems in my thread are automating the hunch.

    Not perfectly, and not everywhere. But the co-scientist does not summarize the literature and hand you a reading list. It proposes mechanisms no one has written down, argues them against itself, and ranks the survivors. Evo-2 does not retrieve a gene; it composes one. Whatever you want to call that, it is happening on the discovery side of Popper’s line, in the territory he declared off-limits to method. The unteachable step is being done by a machine that was, in fact, taught.

    What actually got cheap

    Here is where my thread stops being a highlight reel and starts being an argument, because the interesting question is not whether machines can generate hypotheses. They plainly can. The question is what that does to the rest of science.

    The mathematician Noah Giansiracusa has a name for the pattern, which I take up at more length in my book: carpet bombing. When generation gets cheap, you stop being clever about producing candidates and start producing all of them, then sort. It is how AI does mathematics, throwing enormous numbers of attempts at a problem where checking each one is fast. Hypothesis generation is now carpet bombing pointed at nature. The co-scientist can produce more plausible, well-argued, literature-grounded hypotheses in an afternoon than a lab could dream up in a year.

    And that is exactly where the trouble starts, because a hypothesis is not a proof. It cannot be checked in an afternoon. It has to be checked against the world, and the world runs on its own clock.

    When generation was expensive, the scarce, precious act was having the good idea, and verification, while never easy, was not the binding constraint. A scientist had three hypotheses worth testing and a career to test them in. Reverse that. Now the machine hands you three hundred plausible hypotheses, and the binding constraint is the wet lab, the clinical trial, the telescope time, the years. Generation raced ahead. Verification did not move at all, because verification in science is not a faster model. It is reality, taking as long as reality takes.

    This is the same shape I keep finding everywhere AI touches real work, and I wrote about its purest form in mathematics in an earlier piece. Cheap generation does not remove the bottleneck. It moves it downstream and makes it the whole game.

    The field is already learning this the hard way

    You do not have to take the argument on faith, because the correction is already arriving, and it is arriving in the most useful form: from the people who built the tools.

    Google’s co-scientist reached Nature in 2026, with real wet-lab validation in a handful of biomedical cases. Impressive, and I do not want to wave it away. But when an independent researcher carefully re-implemented the system and ran it hard, the finding was sobering. The pipeline reliably produces hypotheses. Whether it actually improves on the underlying model’s raw guesses was not reproducible from one run to the next, and across dozens of attempts on one disease, not a single one of its generated hypotheses matched the paper’s own headline discoveries. The machine is a fountain of plausible ideas. Plausible is not the same as true, and telling them apart is still the expensive part.

    There is a sharper cautionary tale two years older. In 2023 Google reported that around forty new materials had been discovered and synthesized with the help of one of its AI systems. It was held up as a landmark. Then outside chemists went through the results, and an independent analysis concluded that not one of them was actually a net-new material. The generator worked. The verification, done properly and after the fact by humans, is what separated the discovery from the illusion of one. Every honest account of these systems now carries the same caveat, in the developers’ own words: careful experimental validation, peer review, and independent scrutiny are what turn a generated candidate into knowledge.

    That caveat is not a footnote. It is the job.

    Which part was ever the science

    So let me go back to my thread, and to the line I drew without fully appreciating it.

    I think I was reaching for this, and the clue was in Thaler’s line about defining success. The systems I found most exciting were not the ones that produced the most hypotheses. They were the ones tied to a way of checking, the model that rediscovered gravity and could be tested against known physics, the mathematical work where a conjecture could be pursued to a proof. The ones that unsettled me were the pure generators, magnificent at producing possibilities and silent on which ones were real. The 2020 worry and the 2025 worry are the same worry. A physical law you cannot define success for, and a hypothesis you cannot yet verify, are the same problem.

    Popper drew his line to protect justification. He wanted to say that the logic of science lived in the testing, and that the having-of-ideas, however romantic, was not where rigor lived. The machines have now inverted his world in the most ironic way possible. They have automated the part he thought had no method, the hunch, and in doing so they have made the part he cared about, the checking, more valuable than it has ever been. When hunches were scarce, verification could feel like bookkeeping. Now that hunches are infinite and nearly free, verification is the only thing standing between a lab and a year spent chasing a beautifully argued hypothesis that was never going to be true.

    There is a clue to this in something Andrew Wiles once said, that it is bad to have too good a memory if you want to be a mathematician. It sounds backward until you see what he means. The mathematical gift was never recall. It was compression, the knack for throwing away almost everything and keeping the one idea that organizes the rest, which is roughly how Jürgen Schmidhuber defines insight: a better, shorter way to predict what you have seen. A machine with perfect memory and unlimited generation has exactly the strength Wiles warns against and not yet the one he prizes. It can hold everything. It cannot yet tell what to forget.

    Which returns me to the question I started the thread to answer. Could a machine ever do what Darwin did?

    I think the honest answer is now a qualified yes, and it is qualified in a way I did not expect. A system can hold a mountain of observation and propose the organizing idea. It can make the inductive leap I was so sure was ours alone. But watching it happen, I realized I had misjudged where Darwin’s genius actually sat. The leap to natural selection was extraordinary, but the leap was not the science. The science was in the twenty years, in Darwin knowing which single idea out of the dozens he entertained was worth a life’s defense, and in his relentless testing of it against every objection he could invent. The hunch was the cheaper half but we just could not see that while hunches were rare.

    So the machines got the hunches, and they are welcome to them. What Darwin had that they still do not is the judgment to know which hunch was worth everything, and the patience to spend years finding out if he was wrong. That part did not get automated. It got scarcer, and more valuable, than it was when the ideas were hard to come by.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Can Frontier AI Outdo MBAs?

    Three of the top business schools in the country just tested frontier AI on the analytical work their MBAs are trained to do, and the models scored in the high eighties. If you run a company, that number is coming for you soon, probably in a deck that recommends cutting a layer of analysts. So it is worth being exact about what it measures, because the exact answer is more useful, and more limited, than the headline.

    The paper is BusinessCaseBench, from researchers at Wharton, Carnegie Mellon, and Harvard Business School. They drew 615 questions from real business school cases across eighteen disciplines, from strategy and finance to leadership and ethics. Each model read a case cold and wrote its analysis. A separate grader then compared that analysis against the reference solution the instructor had written, item by item. The models never saw the reference. As far as the model was concerned, it faced an open business question with no answer attached, and it answered well. Under the main metric, Claude Sonnet 4.6 covered about 88 percent of what the instructor’s solution contained, and GPT-5.4 about 87.

    The model was not helped by the answer key. It could not see the rubric, did not use it, and produced its analysis from the case alone, the way a consultant works from a brief. These were, from the model’s side, genuinely open questions. So it is fair to expect that a model which writes strong analyses on 615 unseen cases will write a strong analysis on the 616th, which is your real one.

    The problem is that “the model will write a similar analysis” and “you will get a similar result” are different claims, and only the first one is what the benchmark tested.

    A score is a comparison

    “The model scored 88 percent” is not a fact about the model’s answer by itself. It is a fact about that answer measured against a standard. Two things had to exist for the number to exist: the analysis the model wrote, and the instructor’s solution it was checked against. The score is the relationship between them.

    Now move that setup to your company. The model can still write the analysis. But the standard it gets checked against does not exist. The 88 was a statement about how well the answer matched a known-good answer.

    This is not word games. It is the difference between “the model is competent” and “you can trust the output.” The first is about the answer. The second is about checking the answer, and checking requires a standard. The benchmark supplied the standard. Your hardest decisions do not.

    Knowable-but-hidden is not the same as unknowable

    The tempting reply is that the real world is just the benchmark with the answer hidden. The model handled hidden answers fine; a real decision is one more hidden answer.

    But the benchmark’s answers were not hidden. They were knowable in the first place. The professor had already worked the case. A correct answer existed; the model simply was not shown it. That is a different situation from the one you are in when you decide whether to enter a market or restructure a division. There, no correct answer exists yet. It has not been written by anyone, because the outcome that would settle it is years away, the criteria for “good” are contested by the people in the room, and there is no counterfactual to check the decision against even after the fact.

    A graded case has a knowable answer the model didn’t see. A live strategic decision has an unknowable answer that does not exist to be seen. A student who scores 88 on a past exam she took blind will likely score about 88 on the next past exam. But it does not follow that she will make good venture bets, even though both feel like hard open-ended judgment, because a venture bet has no marking scheme, then or later. The model is the student. The benchmark is the past exam. Your boardroom is the venture bet.

    So the benchmark is strong evidence for a real claim: frontier models are good at producing structured business analysis, and getting better fast. It is not evidence for the claim that the score predicts a good outcome on decisions whose standard has to be invented rather than looked up. Inventing that standard, deciding what a good answer to your actual question would even need to contain, is the judgment.

    That argument stands even if the models are excellent.

    The model is fully right about half the time

    The researchers scored the answers two ways. The headline 88 percent is partial credit: how much of the instructor’s checklist each answer covered. Then they ran a stricter count. On how many questions did the answer satisfy the entire checklist, every item, no gaps? That fell to roughly half.

    On cases where a correct answer was knowable, the leading model produced a complete answer about one time in two. The authors put the point in their own title for that result: these are drafts, not verdicts.

    Now extrapolate honestly. If you carry this model into the wild and expect “similar performance,” you are also carrying the incompleteness. The typical output is strong and missing something at the same time. In the study, a human holding the instructor’s solution caught the missing half. In your firm, if you removed the person who could catch it, the missing half is still missing and nothing catches it.

    Grader didn’t check for what shouldn’t be there

    The grader verifies whether each expected point is present. By construction, it does not scan the answer for confident, invented, or wrong material that sits outside the checklist. A response can hit the expected points and also assert three plausible fabrications and still score well, because nothing in the method is looking for the fabrications.

    In the study that blind spot is harmless, because a grader with the solution ignores the extra material. In your company it is the whole risk. The fabricated line rides along inside a well-organized analysis, unflagged, and the reader who could catch it is the domain expert the strong score seemed to make optional.

    Why “drafts, not verdicts” is the expensive finding

    A draft that is 88 percent right sounds like a bargain, and sometimes it is. But it moves the work rather than removing it. Obvious garbage you would catch. Verified truth you could trust. The strong, incomplete, possibly-embellished draft is the expensive case, because it earns your trust on the parts you can see and hides its gaps in the part you would have to already know the answer to find. Catching what it left out requires someone who knows what a complete answer contains, which is exactly the expertise the draft appeared to retire.

    So the generation got cheap and the checking did not, and on this kind of work the checking cannot be sampled down. You cannot review only the flawed answers, because the flawed ones are the ones that look fine. You review all of them, or you review none and call it oversight. An organization that reads the 88 as license to remove the reviewers has not automated the analysis. It has removed its own ability to tell when the analysis is wrong.

    What it means if you run something

    Read correctly, the benchmark is good news. Frontier models are genuinely strong at producing structured business analysis, and that is real leverage you should use. The error is reading a score-against-a-known-standard as a readiness-to-deploy-without-a-standard.

    Two questions to ask of any AI you are about to trust with judgment work.

    Does this task come with a knowable answer, or is defining the answer the actual job? Where a defensible right answer exists, straightforward analysis with a standard you could write down, let the model run, and expect it to perform in the wild about as it did on the bench. That extrapolation is fair, and worth taking. Where the standard itself has to be invented and argued, the model is drafting, and a person still owns the decision.

    And who checks the drafts, and do they know enough to catch what a confident draft leaves out? If the answer is nobody, or nobody who could tell, the high score is measuring something you will not actually get.

    The models are good at the framed question. Your hardest problems arrive unframed. The competitive advantage was never in answering the case. It was in knowing which case you were actually in.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Early Signs of Rehiring

    A short update to an earlier post on the AI jobs debate.

    A few weeks ago I argued that the AI jobs debate was asking the wrong question. My point was that the layoffs of the past year were mostly a bet, not a result. Companies were cutting in anticipation of what AI would let them do, ahead of the implementation that would justify the cut. Cut first, capture the savings later, and hope the two line up.

    The Wall Street Journal reported this week that, for a lot of large employers, they did not line up.

    What the article says

    Big companies are hiring again. The Journal reports that employers from CSX to Alphabet told investors in recent days that they plan to add people, a reversal after eighteen months of treating hiring as a last resort. The CEO of the HR platform Lattice said many companies stopped hiring junior staff on the assumption that AI agents would cover the work, then realized humans are still needed to work alongside the tools. Her line: having coding agents does not mean you stop hiring engineers, and AI sales agents still need salespeople.

    Separately, initial jobless claims fell to 187,000 for the week ending July 18, the lowest level since September 1969. I checked that against the Labor Department figures, and it holds across Bloomberg, Reuters, and CNN. The year opened with predictions of an AI jobs apocalypse. It is currently producing the fewest unemployment filings in nearly sixty years.

    Why this fits my argument

    What the reporting confirms is the narrow prediction. Cutting on anticipation, ahead of what the technology could actually deliver, was premature, and some of it is now being unwound. The rehiring is the anticipation bet reversing. When the country’s largest employers cut on the theory that agents would absorb the work, then hire back because the agents did not, that is a story about companies acting on a capability that was not there yet.

    What the reporting does not confirm is my deeper claim. My argument was that the real constraint is coordination, the cost of making capable systems work together and with the people around them. The Journal says companies are hiring because of cost, limitation, and uncertainty. It does not say they are hiring because their AI deployments fell apart at the seams. So take this as evidence that the premature-cut prediction was right, and as an open question on the coordination claim, which is the one still worth watching.

    Two caveats on the numbers. Economists describe this as a low-hire, low-fire market: layoffs are low, but hiring is soft too, and June’s dip in the unemployment rate owed partly to a shrinking workforce rather than a boom. And a single week of claims data is noisy, prone to summer seasonal swings.

    The line that spoke the truth

    The MIT labor economist Paul Osterman gave the Journal the most honest sentence in the piece. Asked whether companies need more people or fewer, he said no one has any idea. That uncertainty is the actual state of things.

    The jobs question was never how many humans the machines replace. It was whether an organization understands its own work well enough to know what it can safely hand off. Several did not, so they guessed, and some are now hiring back the people they let go. That is a coherence story, and it is the one I will keep following.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • The Moat Is Coherence

    For most of my career, the systems I built were valuable because they were hard to build.

    At research labs and then inside large enterprises, my teams and I built machine learning systems that took months of work, rare expertise, real budgets, and long struggles for resources. The difficulty was not incidental to the value. It was the value. If a competitor wanted the same capability, they needed the same scarce people, the same time, and the same money. Scarcity of execution was the moat. We just never had to call it that, because it had always been true.

    Looking back now, I feel two things at once. Pride, because the work was genuinely good. And vertigo, because a lot of it could be stood up today in a weekend, by a small team with a subscription. The capability I spent years of my life building is being commoditized.

    Here is the part that took me longer to see. The part of that work that mattered most was never the part that was hard to build. It was knowing what to build, what the data actually meant, which requests to push back on, and which impressive system would quietly make things worse. That part has not commoditized. That part is what this piece is about, because the world’s highest-grossing law firm recently bet half a billion dollars on it.

    The sentence every boardroom should read

    Kirkland & Ellis, which booked $10.6 billion in revenue in 2025, recently announced it would spend roughly $500 million over the next few years building its own AI platform rather than relying only on the tools its competitors can buy. Its chairman, Jon Ballis, compressed the logic into one line. Widely available AI tools, he said, are “raising the floor for everyone.” But, he added, “we don’t get hired for the floor.”

    That is one of the most clarifying things a business leader has said about AI strategy in two years, and it cuts against how most companies are spending. The prevailing assumption is that advantage comes from having the most capable AI. Buy the best models, deploy the most agents, automate fastest, and you win. This was a reasonable playbook for almost every previous technology. It is wrong about this one.

    We have run this experiment before

    It is wrong because capability is commoditizing, and we know what happens when capability commoditizes, because it has happened to every general-purpose technology in modern history. Electricity was an advantage for the firms that could generate it, until it became a utility and the advantage migrated to what firms did with it. Computing was an advantage for firms with mainframe access, until computing became accessible and the advantage migrated to data and process design.

    Nicholas Carr made himself famous, and briefly infamous, by calling this pattern in 2003. His Harvard Business Review essay “IT Doesn’t Matter” argued that information technology was following electricity and the railroads from proprietary advantage into shared infrastructure. The argument was bitterly contested at the time. Two decades later it looks prescient. The firms that built durable advantage in the computing era were not the ones with better hardware. They were the ones that developed distinctive organizational capabilities for using what everyone could buy.

    Michael Porter gave us the vocabulary for why. Operational effectiveness, doing the same things better, is necessary but not sufficient, because best practices diffuse. Strategy is doing things rivals cannot easily match. Tools that a thousand firms can license are operational effectiveness by definition. They raise everyone’s floor at once, and a tool that raises every floor confers advantage on none of them. That the operating model is the moat, more than the model, is a case I made in detail recently, building on McKinsey’s own evidence. If your AI is the same as the AI across the negotiating table, you have spent money to keep pace, not to pull ahead. Ballis’s floor and ceiling is Carr’s argument and Porter’s distinction, restated by a customer with $500 million on the table.

    Even the people selling the technology concede the mechanism. Larry Ellison, whose company is staking its future on enterprise AI, argues that because every major model trains on the same public internet, model outputs are converging and differentiation is eroding. He is right. Shared inputs produce shared reasoning. The implication the vendors are less eager to draw is that buying more of a commoditizing capability is not a strategy. It is a subscription.

    Where Ellison’s answer falls short

    Ellison’s proposed moat is proprietary data. That is closer, but it imports a mistake. Every enterprise already has data, most of it fragmented across systems, contradictory between departments, disconnected from the outcomes it produced, and ungoverned. Pour that into a powerful reasoning engine and you do not get insight. You get fast, confident reasoning over an incoherent picture, which is worse than slow reasoning, because the confidence hides the incoherence. Having data is not scarce.

    What is scarce is data made coherent: owned, reconciled across the organization, connected to the outcomes it produced, and trustworthy enough to reason over, joined to the judgment of the people who know what the numbers mean. I argued earlier in this series why no vendor can sell you this. The platform that stores and retrieves your data is plumbing, and plumbing commoditizes. The coherence is the asset, and it compounds.

    What Kirkland is actually buying

    Read the Kirkland decision through that lens and it stops looking like a technology splurge and starts looking like strategy. The firm is not paying $500 million for data it already owns or software it could license. Anyone can download its public filings. What cannot be downloaded is how the firm decides: the judgment of 250 of its lawyers, 100 of them partners, encoded in a form every lawyer can draw on for every matter. Kirkland is spending to make its institutional judgment coherent, and to keep it exclusive.

    The tell is in the terms. The outside firms building the platform are barred from selling it to any other law firm. If the value were the technology, exclusivity would not matter. Kirkland insists on it because it believes the technology is commoditized and the value is the distinctive judgment the system encodes. A shared tool would dissolve exactly that. A shared tool is also a conduit. Every standard and correction a firm feeds into it can improve the version its rivals rent tomorrow, which is the leakage I traced in Your AI Usage Exhaust Is Someone Else’s Moat. Read that way, Kirkland’s exclusivity clause is that essay’s prescription written into a contract: close the loop, and keep what you encode inside your own walls.

    Since the May announcement, the pattern has only hardened. Through June, Kirkland added two more exclusive builds, one for private-equity fund formation and one for litigation, and framed each the way it framed the platform itself: a way to capture the firm’s own judgment and knowledge and keep it exclusive to Kirkland. Three deals in roughly five weeks, and the constant across all of them is the insistence on owning what the tools encode.

    None of this means Kirkland is certain to be right. Building rather than buying is a real bet, and it is possible that within a few years a purchasable platform, fed a firm’s own data and tuned to its standards, delivers most of the advantage at a fraction of the cost. Every executive now faces a version of that question. But notice what the question is actually about. It is not whether to have AI. It is what you are trying to own. Kirkland has decided the thing worth owning is not the model and not the data but the coherence that turns both into judgment competitors cannot replicate.

    The wrong scoreboard

    The firms that misread this will keep score by the wrong numbers. They will count agents deployed and measure speed of adoption, and they will mistake a rising floor for a rising position. I understand the pull of that scoreboard better than most, because I spent years on the other side of it, building the hard things the scoreboard rewarded. The hard things are cheap now.

    The firms that read it correctly will ask a harder question: when the models are a commodity and the data is everywhere, what does our organization understand about itself that no competitor can buy?

    That was never the intelligence. It is the coherence of the organization putting intelligence to work. Kirkland just put half a billion dollars behind that proposition. The rest of the market is still buying the floor.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • McKinsey Is Right About the Moat. Here’s the Half It Misses.

    McKinsey published a piece this month that gets the hardest part right. Its argument, in one line: the advantage in AI is not the tools, it is the operating model, and the operating model is the one thing a competitor cannot buy. I agree with almost all of it. I want to push on the part it leaves out, because that part is where most companies are about to get hurt.

    Start with what McKinsey gets right, because it is a lot. Efficiency gains from AI, they argue, will become table stakes as the technology spreads. Operating models, unlike software, cannot be purchased or copied overnight, so the durable moat is the organization, not the model. They have the numbers to go with it. Only about a fifth of companies have fundamentally redesigned how they work around AI. Top performers are three times more likely to have done that redesign, and twice as likely to redesign the workflow before choosing the tool. And AI programs run as technology projects fail at more than an eighty percent rate, because they optimize the tool instead of changing how the company works.

    If you have read anything I have written, you know why I would agree. This is the commoditization argument. Capability is becoming universal, so capability stops being the edge, and the advantage moves to what the organization can do that a competitor cannot copy. McKinsey and I are looking at the same shift.

    Here is where we part.

    The scissors cut both ways

    McKinsey’s best image is what they call the complexity scissors. As a company grows, revenue grows not on a line but a curve – fast initially and flattening later. But coordination costs, the meetings and committees and management layers, keep climbing. Plot the two lines and they open like a pair of scissors. The gap between them is why big companies slow down, and why fewer than one in ten sustain returns above their cost of capital over a decade.

    Their prescription is to use AI to close the scissors. Route decisions through a central orchestration layer. Hand coordination work to agents. Flatten the org. Get faster.

    This is the step I want to stop on. AI can close the scissors. It can also open them wider, and nothing in the redesign itself tells you which one you are going to get.

    McKinsey is right that agents let you route around the old coordination layer. What they underplay is that agents build a new coordination surface underneath, invisible, machine-speed, and owned by no one.

    Every autonomous system you add to speed up a workflow is also a new thing that has to stay consistent with every other autonomous system. A support agent and a billing agent that make different assumptions about the same customer have not reduced coordination cost. They have created a new kind of it, one that does not show up in a meeting because no human is in the loop to notice. McKinsey is right that agents let you route around the old coordination layer. What they underplay is that agents build a new coordination surface underneath, invisible, machine-speed, and owned by no one. Their own report admits the danger in a single line: when agents run inside workflows that were not redesigned for them, errors propagate across the company at machine speed. That is the coordination trap, and it is produced by the very rewiring they recommend.

    So the redesign is not the safe move and the caution. The redesign is the risk. Done with the discipline to keep the new systems coherent, it closes the scissors. Done as a race to orchestrate and flatten, it opens them, and it does so faster than the old human version ever could, because now the coordination failures happen at the speed of software.

    This is not just my read of the mechanism. IBM’s 2026 study of two thousand CIOs and CTOs found the same trap from the inside. Companies that chase speed let business units move ahead while governance falls behind, gaining local velocity and losing containment. Companies that chase safety slow deployment under review until oversight becomes unmanageable. Both paths, in the study’s own words, accumulate strategic debt. That is the point. The rewiring does not have a safe default. It has two ways to fail and one narrow way to work, and the narrow way runs through coherence.

    The rewiring does not have a safe default. It has two ways to fail and one narrow way to work, and the narrow way runs through coherence.

    Their own examples make the point

    Look closely at the cases McKinsey uses, because they prove the thing the article does not quite say out loud.

    The copper miner they profile did not win by deploying more AI. It won by building modular models where roughly sixty percent of the code from the first site was reusable across the next six, so each deployment got faster and cleaner than the last. That is a coherence story wearing a productivity headline. The reuse is only possible because someone designed the systems to fit together before scaling them. The car maker they profile shrank a planning team by more than eighty percent, but the win was not the headcount. It was that the coordination layers between data and decision compressed, because the workflow was redesigned as one coherent thing instead of a chain of handoffs.

    In both cases the value came from making the systems cohere, not from the number of systems shipped. McKinsey files this under operating-model redesign. I would file it more precisely: the redesign worked because it was coherent, and it would have failed if it were not. The article treats coherence as a happy property of good redesign. I think it is the whole variable, and that a redesign optimized for speed without it produces the opposite result in the same enterprise.

    What to actually do differently

    If you take McKinsey’s advice and only McKinsey’s advice, you will redesign for speed and measure yourself on how fast you moved. That is the eighty-percent-failure path wearing better clothes, because deployment speed is exactly the vanity metric that hides the debt building underneath.

    The addition is small to state and hard to do. Before you rewire a workflow around agents, decide how those agents will stay consistent with the rest of the company as they multiply. Build the ability to see what your autonomous systems are doing in aggregate, contain them so a failure in one stays in one, and set in advance how much each is allowed to decide. Then measure the redesign not by how fast it shipped but by whether the organization got more coherent or less as it grew. A team that retired four brittle systems and shipped nothing new may have improved your position more than the team that shipped fourteen agents into the trap.

    McKinsey is right that the operating model is the moat, and right that most companies are getting this wrong by treating AI as a tool to buy rather than a business to redesign. The correction I would add is that the redesign has a failure mode of its own, and it is the one nobody is watching for. The winners will not be the companies that rewire fastest. They will be the ones that rewire coherently, which is a slower thing to say and a harder thing to build, and the only version that closes the scissors instead of opening them.

    The winners will not be the companies that rewire fastest. They will be the ones that rewire coherently.

    That is the subject of my book, Coherence, and of everything I am writing here between now and launch. If the rewiring is on your desk right now, the one-page tool I use to sort what to automate, what to redesign, and what to leave alone is the first thing I send when you join the list at coherise.com.