Author: Madhu Shashanka

  • Coding Got Easy, But What Kind?

    A programmer named Senko Rašić published an angry post this week. He is angry at a slogan going around: “code was never the hard part.” He calls it an insult to programmers. The post hit Hacker News and Lobsters and drew more than five hundred comments.

    Reading the comments, I noticed one thing over and over. They argue about what the word “coding” means. Some people say it is the easy part and always was. Others say it is the whole job and always was. They are not disagreeing about difficulty. They are using one word for two different things: producing code, and building something that holds.

    I have a stake in this. I write that execution has gotten cheap and coherence is the hard thing now. Skim that fast and you might file me under the same slogan.

    The hard part was always building well

    Producing code that runs is one thing. Building something that holds is another: the design that survives real data, the structure that stays coherent as the system grows, the choices that still make sense a year later when the author is gone. That second thing was always the hard part, and it was always the job.

    Before AI, doing it well was expensive, because it took a skilled person and their time. Doing it badly was possible, but it was slow and the result was bad. Nothing about producing software was cheap.

    Cheap is what AI added. A model writes a working function or a working page in the time it takes to describe it. Someone with no training can now get running code out of a sentence. That is new, and I will not soften it. Execution got cheap.

    Cheap is not the same as bad

    AI writes good code in places, and the commenters who said so are right. The quality is uneven, and the unevenness has a shape. The models are strong where they had the most to imitate, the patterns written out in public a million times: CRUD, forms, glue code, the endpoint that reads the database and returns JSON. They are weaker where the examples run out, on scientific computing, embedded work, anything performance-critical. One commenter put it well: AI kills it on problems with a thousand forum posts, and you do not point it at the ten-billion-dollar machine headed to Mars.

    There is a reason coding is where AI advanced fastest. Code is checkable. It runs or it does not, the tests pass or they do not, and my book argues that AI improves fastest exactly where the work can be checked. The checkable parts of building will keep getting cheaper and better. What stays hard is the judgment no test can catch.

    That was never the job

    I wrote recently about a delete button from a data engineering thread. An engineer built a button to remove a record. It took the record off the screen. It left behind the three related records the original had created when it was made. The screen looked right. The data underneath was orphaned.

    Producing that button is the cheap part. A model does it now. Knowing it had to clean up three records you cannot see is the judgment. That was the hard part, and that was the job.

    The comparison people keep reaching for

    Watching the threads, I noticed people reaching for the same analogy – writing. Anyone can put words into sentences that flow and reach a point. A model does it fluently. That was never what made someone a writer. What makes a writer is the choice of words and their order. The same point can land or die on those choices. One commenter compared it to a novel, where clean sentences were never the hard part.

    AI did not invent bad prose but made fluent-looking prose free and endless. “Code was never the hard part” runs the same move as saying words were never the hard part of writing. About spelling, it is a shrug. About writing, it is an insult. The slogan gets both out of the same words by letting you hear the first while it means the second.

    The swap, named

    The slogan is true about producing code and false about building well. It earns its credibility on the first and spends it on the second, where the conclusion is that coders are now optional. One commenter worried the line would harden into a truism for business leaders. That is the reader I have in mind.

    If someone read me as saying coders no longer matter, they would be making the same swap. The thing I say got cheap is producing code. The thing I call hard is building something that holds. Those did not both get cheaper. Anyone worried about systems built fast and badly is saying that building them well is still hard. You cannot write about that problem and also believe the work is trivial.

    The right version of the slogan is my argument

    Some people say the line and do mean something reasonable. On Lobsters, a commenter called lcamtuf said Senko was reading it uncharitably, and that the real meaning is that producing lines of code was never the bottleneck. That is true. If coding was fifteen percent of an engineer’s week, automating all of it buys back about a sixth of the week. You aimed the speedup at the fastest part of the job. Another commenter reached for Amdahl’s Law to make the same point with a number.

    Then lcamtuf added the part that matters most to me. Some friction, he said, was good, because it stopped people from building software that was unnecessary or unmaintainable. That friction was a filter. It did not stop people from producing code. It stopped building-without-judgment from becoming load-bearing, because getting anything shipped had to pass through people whose time was scarce. That filter is the thing AI removed. The difficulty of building well did not leave. It stopped being enforced.

    The same thing, at two heights

    Building well has a name once you zoom out. Making each part fit a whole you cannot see all of at once is coherence.

    The delete button is a coherence failure inside one function. The record fit the screen and broke the data underneath. Agentic slop is the same failure across a company. Each workflow works on its own, and the enterprise stops making sense. Senko is defending coherence in one program. I am after it across an organization.

    The slogan is backwards. The hard part of code was never the typing. It was always the judgment, and here is the part that lasts. Where a design was already worked out a thousand times in public, the model has something to copy, and it copies well. Where the design is new to your system, there is nothing to copy and no test to guide it. That judgment stays human, because it is both uncheckable and unwritten.

    Producing code got cheap. That judgment never did. It used to get paid up front, in salaries and review and the time of people who knew what they were doing, where you could see the cost. Now it shows up late, in the corners, with no owner.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Uninformed Expectations, Five Years Later

    In January 2021 an interviewer asked me for the biggest roadblock to AI adoption. My answer came down to one thing: expectations. I called them uninformed. Organizations believed AI was a drop-in component that would improve whatever process it touched. That belief, mixed with a fear of missing out, pushed companies to rush. I wrote that they were reaching for “a new shiny hammer looking for nails,” and that disillusionment would follow.

    I still think that was right. The disillusionment arrived on schedule. But I was also wrong.

    Two versions of one mistake

    The 2021 belief was easy to state. Buy the model and the value follows. AI was a part you slotted into an existing workflow. Reality punished that belief fast, because building anything real was hard. You needed engineering time, data science help, and a business case strong enough to justify the spend. Projects that underestimated the work usually stalled before they shipped. The return died at the front of the pipeline, where the building happened.

    The belief has since turned inside out. Building is easy now. A product manager can stand up an agent in an afternoon with no engineering queue in sight. So the new expectation is that easy building means easy value. If a workflow takes a day instead of a quarter, the returns should take care of themselves.

    They don’t. And the reason traces back to the same root as before.

    Capability was never the outcome

    Both beliefs make the same move. They treat capability as if it were the outcome. Capability is what your AI can do. The outcome is a separate thing: what your organization still has once the AI has done it. The space between those two is where the return leaks away.

    In 2021 that space was easy to see, because scarcity kept it visible. To build anything, a team had to win scarce engineering time. Winning it meant convincing people outside the team. A budget owner. An architect. A security reviewer. Nobody designed that as oversight. It was just the price of a scarce resource. It still worked like a filter. Weak ideas died in the queue, and only the ones someone could defend reached production. Scarcity was doing quiet work that never showed up on an org chart.

    That filter is gone. When building costs almost nothing, nothing stops a weak idea from becoming a running system. Fifty teams can each ship their own agent, every one of them green on its own dashboard, and no single person owns the question of what they add up to. The return still leaks. It leaks at the far end of the pipeline now, in the cost of coordinating systems nobody mapped.

    The newest version of the old belief

    What is the belief being sold right now? The frontier model vendors have told the market that the gap to ROI is expertise. You have the models. What you lack is people who know how to wire them into your environment. So the labs send forward deployed engineers. They embed at your site, build the integrations, tune the configurations, debug the odd behavior, and leave a working system behind.

    The role exists because my diagnosis is right. Building the model was never the hard part. Deploying it inside a messy enterprise is. FDEs are a real answer to that, and a good one. They are also an answer to the wrong problem, and their structure guarantees it.

    Start with the incentive. An FDE works for the vendor. Success for them means adoption and a satisfied customer. They have no reason, and usually no mandate, to tell you that a deployment conflicts with a system three departments away that they cannot see, or that it will cost you more in complexity than it returns in efficiency, or that the right call is to not build it. The most valuable act of coherence is sometimes the word no. You cannot buy that from the party paid to say yes.

    Then there is what they can see. An FDE embedded in one business unit knows that unit. They have no view of the other agents running across the company, or the coordination surfaces their new system quietly creates. The failures I worry about do not come from one bad deployment. They come from the sum of many reasonable ones. No FDE is positioned to see the sum.

    And there is what they leave behind. When the engagement ends, the deepest understanding of why the system was built that way, what it assumes, and how to change it safely often leaves with them. You inherit a running system and a dependency, not the knowledge to oversee it over time.

    So the FDE belief is the 2021 belief again, dressed for 2026. In 2021 it was “buy the model and value follows.” Now it is “add the deployment experts and value follows.” Both stop at capability. Deployment velocity is still capability. It is the thing every competitor can rent from the same labs, on the same terms, in the same quarter. It gets the system live. It does not decide whether the system should have gone live at all, and it does not hold the enterprise together once fifty of them are running. FDEs solve half the problem. The half they cannot touch is the one that decides your return.

    Why the old advice still works

    Back then I offered a few tips. Go slow. Pick narrow, well-defined use cases. Get a quick win on the board. Kill a project when the evidence says to, and keep sunk cost from making that call for you. I stand behind every word of it. What surprises me is why it still holds.

    In 2021 that discipline was prudence. Ignore it and scarcity would punish you, so the advice helped you get through the queue with something that worked. Today the same discipline is close to the only filter left in the building. “Go slow” used to save you from a stalled project. Now it is most of what stands between you and a sprawl of systems you cannot see. “Kill the project” used to fight sunk cost. Now it fights the agent that keeps running because nobody confirmed it should stop.

    The advice is unchanged. What used to enforce it for free has disappeared, so the work now falls to you. Scarcity handled a crude version of this by accident. You have to handle the real version on purpose.

    The name I was missing

    When I wrote that post I was reaching for something I could not name. I knew rushing was dangerous. I knew AI was a means to an end. I had no word for the property that separates a company that gets value from one that gets debt.

    The word is coherence. It is the capacity to see what your autonomous systems are doing, judge whether they are doing it well, and correct them when they are not. In 2021, scarcity supplied a crude version of it by accident. In 2026, you build it on purpose or you go without. Once execution gets cheap, coherence becomes the thing that decides whether all that abundance turns into advantage or into cleanup.

    The disillusionment I flagged five years ago is still on its way to a lot of organizations. It reaches them from the opposite direction now. Back then it came from systems that were too hard to build. Now it comes from systems that are too easy to build. The belief underneath is the one I named in 2021. Stop mistaking what AI can do for what your organization will keep, and build the coherence that turns the first into the second.

    The full argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • The Whole Book In Fifteen Sentences

    Newsletter – Edition 3

    A quick milestone. Copy-editing on the book is finished, and it has moved into design. The words are settled. Now it becomes an object you can hold.

    Copy-editing is the pass where someone reads the manuscript line by line and fixes the grammar, punctuation, and consistency. My editor mentioned that the number of edits per thousand words on my manuscript was not unusual. I hope she was being true and not just being kind.

    Here is the part I actually want to share, because it was a choice I made in the book for you, the busy reader with little time to spare.

    At the end of every chapter there is a box called Key Takeaways. It answers six questions, in order. The one thing to remember. The demonstration. Why it holds. How to recognize it. What it changes. Where it goes next.

    The boxes do one more thing together. Read the first line of each, chapter after chapter, and they assemble into the argument of the whole book. Fifteen sentences, front to back. Here it is.


    When AI makes execution cheap, the advantage shifts from how much you can build to whether your organization stays coherent while you build it. The whole book, in order, reads as follows.

    Part 1: The Inversion

    1. Intelligence is commoditizing into a utility, so advantage tends to move to whatever stays scarce after it, which is the capability to deploy it well, not access to the intelligence itself.
    2. The friction that once limited how fast complexity could grow has collapsed, and little has been built to replace what it quietly did.
    3. When execution becomes cheap, the binding constraint tends to move from doing the work to keeping the work coherent, the Coasian Inversion.
    4. Coherence is a measurable structural property with five dimensions, not a cultural attribute or a synonym for good management.

    Part 2: The Failures

    1. Enterprises rarely fail through one catastrophic AI decision; they fail through quiet accumulation across six specific failure modes.
    2. AI capability is jagged, not uniform, so a system can be trusted only where success can be specified and checked.
    3. Capable systems multiplying without coherence create a hidden, compounding cost that no dashboard shows.
    4. Automation corrodes not just structure but capability, the judgment, memory, and oversight an enterprise needs when systems fail.
    5. Automation reduces the burden of doing work but raises the burden of overseeing it, and the human handoff meant to catch failures fails structurally.
    6. Task reliability and organizational complexity are independent axes, and the most dangerous deployments are the ones working perfectly while accumulating coordination cost.

    Part 3: The Discipline

    1. Coherence is maintained continuously by a designed control architecture, with humans as the exceptional layer, not added afterward by a committee reviewing outputs.
    2. Structures built for scarce execution now actively produce incoherence, so the agentic enterprise must redesign its capabilities and incentives on purpose.
    3. Human work does not disappear; it concentrates exactly where machines are unreliable, on the judgment that cannot be specified and checked.
    4. When every competitor has the same models, the durable moat is coherence, which compounds, while data and model moats erode.
    5. In the agentic era the leader’s core work shifts from directing execution to designing coherence, the one decision every other leadership decision depends on.

    When everyone can go fast, going fast is no longer the advantage. What wins is whether the organization can go fast without coming apart.


    That is the spine. The chapters are the muscle around it.

    The week in ideas

    Four posts from the past week.

    AI Still Needs Human Bosses? The New York Times handed an AI agent three office jobs. It wrote clean code in minutes, then failed at judgment. It could not upload a file, so it quietly marked the task done. It read employees on leave as cuttable roles. Firms are thinning the supervisory layer that catches exactly these errors, while installing systems that consume more of it. That reversal is the thesis of the book, and one version of it has already reached a federal courtroom. Weigh in on LinkedIn…

    “The New Normal Because Faster” A viral Reddit thread about an enterprise platform deployment that went badly. I cannot verify a word of it, so I make no claim about any company. But the pattern the accounts describe is the one the book predicts. A delete button that clears the record from the screen and orphans the three hidden records it created. Locally correct, globally broken. A commenter named the whole thing in five words: the new normal because faster. Execution got cheaper. Coherence did not. Weigh in on LinkedIn…

    Safest Car on the Road, Yet Parks in the Fire Lane Waymo is far safer than human drivers across 50 million miles and still collects parking tickets across my hometown of Austin. Two different failures live in that story. Jaggedness, where a system is superhuman at driving and stumped by a handicap spot. And the deeper one, where a firefighter has full authority over the car and no lever to move it. Three hundred cars each parking rationally still block the same church garage. The tickets are the city’s crude, correct instinct: price the incoherence when you cannot redesign the system. Weigh in on LinkedIn…

    Early Signs of Rehiring A short update to an earlier post. Big employers from CSX to Alphabet are hiring again after eighteen months of treating hiring as a last resort. The narrow prediction held. Companies cut on the bet that agents would absorb the work, then hired back when the agents did not. The deeper coordination claim stays an open question. Best line, from an MIT economist asked whether firms need more people or fewer: no one has any idea.

    One thread runs through all four. The machine can produce the output. A human still owns the part with no dashboard: the judgment, the handoff, the curb no single car is responsible for. That is my book in one sentence, which is a convenient thing to be able to say now that the fifteen are sitting above.

    Before you go

    Design is where a manuscript stops being a document and starts being a book. I will share the cover here first when it is ready.

    If the book is why you are here, it is Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. Those stories are how I can learn.

  • Could a Machine Have Had Darwin’s Idea?

    In November 2020 I started a thread on X. I have barely used the platform these past couple of years, and I only came back to this thread recently, by accident, when some of my own old posts surfaced in front of me. Reading them in one sitting was strange, because a thread I had added to piecemeal over five years turned out to have been about one question the whole time.

    It opened with Darwin. Someone I followed had posted, amazed, about what evolution manages to build, and I wrote that what amazed me was something else: the intellectual leap required to infer the theory in the first place, from nothing but empirical observation and years of focused study. I said I was not sure I fully appreciated Darwin’s intellect and dedication. Then, in the next post, I said the thing the whole thread has been chasing ever since. That kind of leap, I wrote, is the kind of intelligence AI should be aiming for, not recognizing cat faces or copying tasks people already do well.

    That was the bar I set in 2020, and it was a high one. Darwin spent decades buried in finches and barnacles and pigeon breeders and fossil beds, a mountain of unconnected observation, and then made a leap: one idea, natural selection, that reorganized all of it at once. That inductive move, from a heap of messy particulars to the principle that explains them, struck me then as the most formidable act an intellect can perform, and the part of real science I was least sure a machine could touch. So I kept a list, adding to it whenever I came across a case of AI looking like it was helping advance science rather than just crunch it, to watch whether anything ever cleared the bar.

    Reading the thread back now, it traces an arc I did not plan, and it worried at the right problem from the start. Within days, in November 2020, I posted a New York Times piece on Max Tegmark’s group, whose neural network had recovered a hundred physics equations from raw data, and I pulled out the catch that the physicists themselves named. Tegmark was candid that the machine could retrieve the formulas but not yet the deep principles beneath them, the quantum uncertainty or the relativity that would explain why the formula holds. And Jesse Thaler, the MIT physicist directing the new AI-and-physics institute, put his finger on why. AI wins at games because a game has a well-defined notion of success. “If we could define what success means for physical laws,” he said, “that would be an incredible breakthrough.” Proposing the theory, in other words, was the hard part, and it was hard precisely because you could not say in advance what would count as getting it right.

    The entries that moved me most in those early days were not about AI at all. They were about humans doing the thing I wanted to see a machine do. Around the same time I was reading about Tibor Gánti, the Hungarian biologist who deduced from first principles what the simplest possible living thing must be: a metabolism, a way to store information, and a membrane, three systems that have to be coupled or the organism dies. I wrote at the time that this was the kind of inductive thinking we should be teaching children. Later I followed a related idea, assembly theory, which proposes a single elegant handle on complexity: count the minimum number of steps needed to build a molecule, and use that number to test for the presence of life on other worlds. A theory reduced to one measurable quantity, aimed at one of the hardest questions there is. It may not survive contact with the evidence; I later noted that a study found some minerals scoring above the threshold the theory sets for life. But right or wrong, it was the move I admired, the leap from observation to a principle sharp enough to be tested.

    Then the machine examples accumulated. In December 2021, a Nature paper where neural networks guided the intuition of mathematicians toward new conjectures in knot theory, the machine surfacing the pattern and the human still making the leap. In March 2022, a system that rediscovered Newton’s law of gravitation from the motion of the planets and wrote it back out as a symbolic equation. Then more symbolic regression, pulling laws from data. Then, in February 2025, the entries that made me sit up: Google’s AI co-scientist, generating and ranking novel research hypotheses on its own, and Evo-2, a model that does not just read genomes but writes them. And I drew a line I did not fully understand at the time. I noted that the most exciting work was different from the systems that merely “generate plausible hypotheses from an extremely large space of possibilities.”

    Five years of watching, and that distinction turned out to be the whole story.

    The step that was supposed to have no method

    For a century, the standard account of science drew a hard line between two acts.

    One is coming up with the idea. The other is checking whether the idea is true. Karl Popper named these the context of discovery and the context of justification, and he was blunt about which one belonged to philosophy. Testing a hypothesis has a logic. Having one does not. In his words, the act of conceiving a theory “neither calls for logical analysis nor is susceptible of it.” The initial leap from a pile of observations to this might be why was, he thought, a matter for psychology, not method. A hunch. The unteachable part.

    That is where the whole romance of science lived, and Darwin’s leap is its patron saint. Kekulé dreaming the benzene ring as a snake biting its tail. Fleming noticing the one culture plate that had gone wrong in an interesting way. Darwin holding twenty years of specimens in his head until they resolved into a single idea. We told these stories because the leap seemed to come from nowhere, and coming from nowhere was the point. You could train someone to run an experiment. You could not train the hunch.

    The systems in my thread are automating the hunch.

    Not perfectly, and not everywhere. But the co-scientist does not summarize the literature and hand you a reading list. It proposes mechanisms no one has written down, argues them against itself, and ranks the survivors. Evo-2 does not retrieve a gene; it composes one. Whatever you want to call that, it is happening on the discovery side of Popper’s line, in the territory he declared off-limits to method. The unteachable step is being done by a machine that was, in fact, taught.

    What actually got cheap

    Here is where my thread stops being a highlight reel and starts being an argument, because the interesting question is not whether machines can generate hypotheses. They plainly can. The question is what that does to the rest of science.

    The mathematician Noah Giansiracusa has a name for the pattern, which I take up at more length in my book: carpet bombing. When generation gets cheap, you stop being clever about producing candidates and start producing all of them, then sort. It is how AI does mathematics, throwing enormous numbers of attempts at a problem where checking each one is fast. Hypothesis generation is now carpet bombing pointed at nature. The co-scientist can produce more plausible, well-argued, literature-grounded hypotheses in an afternoon than a lab could dream up in a year.

    And that is exactly where the trouble starts, because a hypothesis is not a proof. It cannot be checked in an afternoon. It has to be checked against the world, and the world runs on its own clock.

    When generation was expensive, the scarce, precious act was having the good idea, and verification, while never easy, was not the binding constraint. A scientist had three hypotheses worth testing and a career to test them in. Reverse that. Now the machine hands you three hundred plausible hypotheses, and the binding constraint is the wet lab, the clinical trial, the telescope time, the years. Generation raced ahead. Verification did not move at all, because verification in science is not a faster model. It is reality, taking as long as reality takes.

    This is the same shape I keep finding everywhere AI touches real work, and I wrote about its purest form in mathematics in an earlier piece. Cheap generation does not remove the bottleneck. It moves it downstream and makes it the whole game.

    The field is already learning this the hard way

    You do not have to take the argument on faith, because the correction is already arriving, and it is arriving in the most useful form: from the people who built the tools.

    Google’s co-scientist reached Nature in 2026, with real wet-lab validation in a handful of biomedical cases. Impressive, and I do not want to wave it away. But when an independent researcher carefully re-implemented the system and ran it hard, the finding was sobering. The pipeline reliably produces hypotheses. Whether it actually improves on the underlying model’s raw guesses was not reproducible from one run to the next, and across dozens of attempts on one disease, not a single one of its generated hypotheses matched the paper’s own headline discoveries. The machine is a fountain of plausible ideas. Plausible is not the same as true, and telling them apart is still the expensive part.

    There is a sharper cautionary tale two years older. In 2023 Google reported that around forty new materials had been discovered and synthesized with the help of one of its AI systems. It was held up as a landmark. Then outside chemists went through the results, and an independent analysis concluded that not one of them was actually a net-new material. The generator worked. The verification, done properly and after the fact by humans, is what separated the discovery from the illusion of one. Every honest account of these systems now carries the same caveat, in the developers’ own words: careful experimental validation, peer review, and independent scrutiny are what turn a generated candidate into knowledge.

    That caveat is not a footnote. It is the job.

    Which part was ever the science

    So let me go back to my thread, and to the line I drew without fully appreciating it.

    I think I was reaching for this, and the clue was in Thaler’s line about defining success. The systems I found most exciting were not the ones that produced the most hypotheses. They were the ones tied to a way of checking, the model that rediscovered gravity and could be tested against known physics, the mathematical work where a conjecture could be pursued to a proof. The ones that unsettled me were the pure generators, magnificent at producing possibilities and silent on which ones were real. The 2020 worry and the 2025 worry are the same worry. A physical law you cannot define success for, and a hypothesis you cannot yet verify, are the same problem.

    Popper drew his line to protect justification. He wanted to say that the logic of science lived in the testing, and that the having-of-ideas, however romantic, was not where rigor lived. The machines have now inverted his world in the most ironic way possible. They have automated the part he thought had no method, the hunch, and in doing so they have made the part he cared about, the checking, more valuable than it has ever been. When hunches were scarce, verification could feel like bookkeeping. Now that hunches are infinite and nearly free, verification is the only thing standing between a lab and a year spent chasing a beautifully argued hypothesis that was never going to be true.

    There is a clue to this in something Andrew Wiles once said, that it is bad to have too good a memory if you want to be a mathematician. It sounds backward until you see what he means. The mathematical gift was never recall. It was compression, the knack for throwing away almost everything and keeping the one idea that organizes the rest, which is roughly how Jürgen Schmidhuber defines insight: a better, shorter way to predict what you have seen. A machine with perfect memory and unlimited generation has exactly the strength Wiles warns against and not yet the one he prizes. It can hold everything. It cannot yet tell what to forget.

    Which returns me to the question I started the thread to answer. Could a machine ever do what Darwin did?

    I think the honest answer is now a qualified yes, and it is qualified in a way I did not expect. A system can hold a mountain of observation and propose the organizing idea. It can make the inductive leap I was so sure was ours alone. But watching it happen, I realized I had misjudged where Darwin’s genius actually sat. The leap to natural selection was extraordinary, but the leap was not the science. The science was in the twenty years, in Darwin knowing which single idea out of the dozens he entertained was worth a life’s defense, and in his relentless testing of it against every objection he could invent. The hunch was the cheaper half but we just could not see that while hunches were rare.

    So the machines got the hunches, and they are welcome to them. What Darwin had that they still do not is the judgment to know which hunch was worth everything, and the patience to spend years finding out if he was wrong. That part did not get automated. It got scarcer, and more valuable, than it was when the ideas were hard to come by.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Can Frontier AI Outdo MBAs?

    Three of the top business schools in the country just tested frontier AI on the analytical work their MBAs are trained to do, and the models scored in the high eighties. If you run a company, that number is coming for you soon, probably in a deck that recommends cutting a layer of analysts. So it is worth being exact about what it measures, because the exact answer is more useful, and more limited, than the headline.

    The paper is BusinessCaseBench, from researchers at Wharton, Carnegie Mellon, and Harvard Business School. They drew 615 questions from real business school cases across eighteen disciplines, from strategy and finance to leadership and ethics. Each model read a case cold and wrote its analysis. A separate grader then compared that analysis against the reference solution the instructor had written, item by item. The models never saw the reference. As far as the model was concerned, it faced an open business question with no answer attached, and it answered well. Under the main metric, Claude Sonnet 4.6 covered about 88 percent of what the instructor’s solution contained, and GPT-5.4 about 87.

    The model was not helped by the answer key. It could not see the rubric, did not use it, and produced its analysis from the case alone, the way a consultant works from a brief. These were, from the model’s side, genuinely open questions. So it is fair to expect that a model which writes strong analyses on 615 unseen cases will write a strong analysis on the 616th, which is your real one.

    The problem is that “the model will write a similar analysis” and “you will get a similar result” are different claims, and only the first one is what the benchmark tested.

    A score is a comparison

    “The model scored 88 percent” is not a fact about the model’s answer by itself. It is a fact about that answer measured against a standard. Two things had to exist for the number to exist: the analysis the model wrote, and the instructor’s solution it was checked against. The score is the relationship between them.

    Now move that setup to your company. The model can still write the analysis. But the standard it gets checked against does not exist. The 88 was a statement about how well the answer matched a known-good answer.

    This is not word games. It is the difference between “the model is competent” and “you can trust the output.” The first is about the answer. The second is about checking the answer, and checking requires a standard. The benchmark supplied the standard. Your hardest decisions do not.

    Knowable-but-hidden is not the same as unknowable

    The tempting reply is that the real world is just the benchmark with the answer hidden. The model handled hidden answers fine; a real decision is one more hidden answer.

    But the benchmark’s answers were not hidden. They were knowable in the first place. The professor had already worked the case. A correct answer existed; the model simply was not shown it. That is a different situation from the one you are in when you decide whether to enter a market or restructure a division. There, no correct answer exists yet. It has not been written by anyone, because the outcome that would settle it is years away, the criteria for “good” are contested by the people in the room, and there is no counterfactual to check the decision against even after the fact.

    A graded case has a knowable answer the model didn’t see. A live strategic decision has an unknowable answer that does not exist to be seen. A student who scores 88 on a past exam she took blind will likely score about 88 on the next past exam. But it does not follow that she will make good venture bets, even though both feel like hard open-ended judgment, because a venture bet has no marking scheme, then or later. The model is the student. The benchmark is the past exam. Your boardroom is the venture bet.

    So the benchmark is strong evidence for a real claim: frontier models are good at producing structured business analysis, and getting better fast. It is not evidence for the claim that the score predicts a good outcome on decisions whose standard has to be invented rather than looked up. Inventing that standard, deciding what a good answer to your actual question would even need to contain, is the judgment.

    That argument stands even if the models are excellent.

    The model is fully right about half the time

    The researchers scored the answers two ways. The headline 88 percent is partial credit: how much of the instructor’s checklist each answer covered. Then they ran a stricter count. On how many questions did the answer satisfy the entire checklist, every item, no gaps? That fell to roughly half.

    On cases where a correct answer was knowable, the leading model produced a complete answer about one time in two. The authors put the point in their own title for that result: these are drafts, not verdicts.

    Now extrapolate honestly. If you carry this model into the wild and expect “similar performance,” you are also carrying the incompleteness. The typical output is strong and missing something at the same time. In the study, a human holding the instructor’s solution caught the missing half. In your firm, if you removed the person who could catch it, the missing half is still missing and nothing catches it.

    Grader didn’t check for what shouldn’t be there

    The grader verifies whether each expected point is present. By construction, it does not scan the answer for confident, invented, or wrong material that sits outside the checklist. A response can hit the expected points and also assert three plausible fabrications and still score well, because nothing in the method is looking for the fabrications.

    In the study that blind spot is harmless, because a grader with the solution ignores the extra material. In your company it is the whole risk. The fabricated line rides along inside a well-organized analysis, unflagged, and the reader who could catch it is the domain expert the strong score seemed to make optional.

    Why “drafts, not verdicts” is the expensive finding

    A draft that is 88 percent right sounds like a bargain, and sometimes it is. But it moves the work rather than removing it. Obvious garbage you would catch. Verified truth you could trust. The strong, incomplete, possibly-embellished draft is the expensive case, because it earns your trust on the parts you can see and hides its gaps in the part you would have to already know the answer to find. Catching what it left out requires someone who knows what a complete answer contains, which is exactly the expertise the draft appeared to retire.

    So the generation got cheap and the checking did not, and on this kind of work the checking cannot be sampled down. You cannot review only the flawed answers, because the flawed ones are the ones that look fine. You review all of them, or you review none and call it oversight. An organization that reads the 88 as license to remove the reviewers has not automated the analysis. It has removed its own ability to tell when the analysis is wrong.

    What it means if you run something

    Read correctly, the benchmark is good news. Frontier models are genuinely strong at producing structured business analysis, and that is real leverage you should use. The error is reading a score-against-a-known-standard as a readiness-to-deploy-without-a-standard.

    Two questions to ask of any AI you are about to trust with judgment work.

    Does this task come with a knowable answer, or is defining the answer the actual job? Where a defensible right answer exists, straightforward analysis with a standard you could write down, let the model run, and expect it to perform in the wild about as it did on the bench. That extrapolation is fair, and worth taking. Where the standard itself has to be invented and argued, the model is drafting, and a person still owns the decision.

    And who checks the drafts, and do they know enough to catch what a confident draft leaves out? If the answer is nobody, or nobody who could tell, the high score is measuring something you will not actually get.

    The models are good at the framed question. Your hardest problems arrive unframed. The competitive advantage was never in answering the case. It was in knowing which case you were actually in.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Sensing Is More Than Measurement

    The Financial Times reported this week on an internal Amazon presentation held July 28. Engineers walked staff through a set of AI cost overruns. The largest used Anthropic’s Claude Sonnet to match author details against product listings on Amazon’s retail site. It ran 860 percent over budget, cost $1.8 million, and never shipped. Engineers said mistakes that were once trivially cheap had become “catastrophically expensive”.

    A financial auditing tool ran about $541,000 over. A logistics project meant to reduce delivery times ran about $134,000 over. Roughly $2.5 million across the three.

    News coverage led with the money. Several outlets called it a coding task. Matching author details to listings is data reconciliation, the kind of work nobody watches.

    The project ran for five months before anyone caught it.

    The data was there

    A senior Amazon employee told the FT that it is difficult to figure out how much anything AI-related costs. Said by someone at the company that runs the cloud everyone else buys AI on.

    Token spend is metered, priced publicly, and billed monthly. Every token that project consumed appeared on an invoice. Engineers attributed the overruns partly to the shift from flat subscriptions to token-based billing, where costs climb whenever a task generates more activity than expected.

    Few things inside a large enterprise are more thoroughly instrumented than a cloud bill. Amazon had the numbers for five months and stayed unaware of them.

    Sensing takes more than measurement. Something has to compare the number against an expectation, notice the gap, and route it to someone who can act while acting is still cheap. Amazon had the number. The comparison and the route were missing.

    No dashboard would have closed this. Someone had to decide that aggregate token spend against declared intent is a thing the company watches, and then own the watching.

    March incidents

    Amazon’s retail website took four high-severity incidents in a single week in early March, including a six-hour failure that locked customers out of checkout, account information, and pricing.

    An internal document prepared for the review meeting identified GenAI-assisted changes as a factor in a pattern of incidents going back to Q3. That reference was deleted before the meeting, according to the FT, which saw both versions. Amazon disputed the reporting and said only one incident involved AI directly, with the root cause an engineer acting on inaccurate advice an AI agent had inferred from an outdated internal wiki.

    Amazon’s response was a 90-day code safety reset across 335 critical retail systems and mandatory senior-engineer sign-off on AI-assisted code from junior and mid-level engineers.

    Work backward from July. Five months of undetected spending starts around February or March. I cannot confirm the detection date, so treat the overlap as inference. Even without it, the shape holds. Amazon added review gates on AI-assisted changes to critical systems while a cost failure accumulated invisibly on a job nobody would call critical.

    Constraint depends on detection. You cannot cap, contain, or price what you cannot see. Amazon reached for the second without the first, which happens because approval steps are visible to leadership and instrumentation is not.

    KiroRank

    Amazon ran an internal leaderboard called KiroRank that ranked employees by AI usage. Staff responded with what they called tokenmaxxing, deliberately inflating consumption to climb the rankings. Amazon discontinued it.

    Every deployment a team builds imposes cost on everyone else. Another surface to watch, another dependency to reconcile, another system someone who did not build it has to understand. The team keeps the benefit while the organization bears the cost. The remedy is to price that burden back to the team creating it.

    Amazon built a price signal pointing the wrong way. Ranking people by consumption pays them to consume. An internal metric carrying status and no cost gets gamed, and this one did.

    The presentation’s own recommendations now include avoiding leaderboards that reward token consumption, and checking whether higher token usage produces useful output.

    The objection

    Amazon frames these as isolated examples of teams learning from one another, and says cherry-picking them does not reflect how teams across the company use AI. For a company its size, a seven-figure surprise is a rounding error.

    The reporting also lacks a base rate. Nobody has said how many AI projects came in on budget for every one that blew up.

    Five months of invisibility still belongs to the control architecture rather than the budget. The same architecture at a company with a $12 million annual AI budget produces the same five months and a different outcome. Amazon can absorb what it cannot see.

    What it costs to fix

    Spending on sensing is bounded and knowable in advance. You can price the instrumentation, the ownership, and the review cadence before committing. The incoherence it prevents accrues silently and surfaces only once addressing it is no longer optional.

    Amazon paid roughly $2.5 million across three disclosed projects. Other companies will meet the same failure without the revenue to absorb it.

    The argument here about sensing, constraint, and priced externalities runs through my book, Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Why Ford Rehired

    Ford has added more than 350 experienced engineers over the past three years after its automated quality systems failed to deliver the results the company expected. Inside Ford they are called gray beards. Some are former Ford employees. Others came from suppliers.

    Ford says it added the specialists “through internal promotions or new talent” to work alongside newer team members. Headlines have called it rehiring people AI replaced but Ford has not used that word, and the reporting does not establish that these specific roles were cut.

    Charles Poon, Ford’s vice president of vehicle hardware engineering, told reporters that AI is a fantastic tool and only as good as the information used to train it. He was more direct about the error: “Mistakenly, we thought that by just introducing artificial intelligence and ingesting the design requirements that we had, that would produce a high-quality product.”

    Poon also said Ford had not paid enough attention in prior years to the experience of its most knowledgeable engineers, the ones who had been through many product cycles.

    Kumar Galhotra, Ford’s chief operating officer, said the company had been relying more and more on automated quality systems before recognizing the approach was not working.

    Ford then ranked first among mainstream brands in the 2026 J.D. Power U.S. Initial Quality Study, its first time since 2010.

    Ford’s explanation

    Ford describes a training data problem. Poon’s version: enhancing the automation and machine learning tools required making sure they were trained by the most experienced individuals.

    Part of that holds up plainly. The specialists do reprogram the AI tools that fell short.

    The jobs

    A missing corpus has a fix with an end date. Sit the veterans down, extract the failure modes they carry, feed the models, thank them.

    But Ford built something with no end date. The specialists run mandatory meetings on quality concerns and hunt for failure points before a part reaches the plant floor. They conduct regular design reviews to identify problems before vehicles reach production. They train junior staff. Galhotra put the shift as moving from a find-and-fix mentality to preventing issues before they occur.

    A standing design review is a verification loop the organization has decided to keep running.

    Where the corpus story runs out

    Tacit knowledge can be captured up to a point. What a senior engineer knows about how a joint fails under a particular thermal cycle can be written down, and should be.

    Judgment applied to an unanticipated case cannot. A reviewer looks at a novel configuration and says it will not hold, for reasons that emerge from the thing in front of them. Enumerate those cases ahead of time and you would have automated the review already.

    The corpus framing implies a completion state, where enough capture makes the humans optional. Ford’s remedy points elsewhere. Mandatory and recurring is what you build once you have concluded the checking does not stop.

    Automating a quality inspection function means automating a verifier, which removes the capacity to tell whether the automation works.

    Ford’s specialists hold the ability to tell the machine it is wrong.

    The order

    Poon’s admission about prior years is the sharpest thing either executive said. Ford’s assumption was reasonable and its sequence was wrong. Capture the expertise, then automate, and the program is defensible. Automate on the assumption that design requirements are sufficient, and you spend three years buying judgment back from suppliers and internal promotions.

    Each step is the precondition for the next. Skipping one relocates its cost to a later point, larger, with fewer options available. Ford turned a knowledge problem into a three-year staffing program, and that program ran alongside more than $1 billion in expected warranty and material costs this year and a quality reputation to repair.

    Mentorship

    Ford is explicit that mentorship is part of the assignment. The specialists work alongside newer team members and train junior staff who never absorbed the institutional knowledge.

    This is how senior judgment gets built. Less experienced people work real problems while someone holding the judgment watches and corrects. The routine cases are the training ground, and automation takes them first.

    An organization that automates routine work and then loses its seniors breaks the pipeline at both ends. Nobody holds the judgment and nobody acquires it. Ford is paying to rebuild both.

    Headcount models calculate savings against the cost of the people. The judgment training pipeline appears nowhere in the model.

    Ford is still deploying AI

    On an autumn 2025 earnings call, Galhotra said Ford was systemically deploying AI across the entire industrial system, including 900 AI-powered cameras across its plants to detect quality issues at the source. Ford kept the cameras. Jim Farley told Bloomberg TV that Ford has AI tools for vision systems, and that most of it comes down to team members paying attention to small details.

    Ford topped the mainstream J.D. Power rankings with the automated systems still running. Experienced people now sit between the systems and the product.

    Experienced engineers sat on Ford’s books as execution capacity, a cost line. Their function was judgment, which is what makes execution capacity safe to deploy.

    Ford is not unusual in getting the order wrong. CNBC reported Robert Half data showing 32 percent of U.S. hiring managers eliminated a role primarily because of AI and later rehired for the same or a similar position. Robert Half’s own summary puts it as more than 3 in 10, and notes the two most common reasons given: the role required institutional knowledge or context AI could not replace, and it involved relationship management AI could not replicate.

    The arguments here about judgment as the layer that cannot be purchased, and about sequence in organizational automation, run through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Early Signs of Rehiring

    A short update to an earlier post on the AI jobs debate.

    A few weeks ago I argued that the AI jobs debate was asking the wrong question. My point was that the layoffs of the past year were mostly a bet, not a result. Companies were cutting in anticipation of what AI would let them do, ahead of the implementation that would justify the cut. Cut first, capture the savings later, and hope the two line up.

    The Wall Street Journal reported this week that, for a lot of large employers, they did not line up.

    What the article says

    Big companies are hiring again. The Journal reports that employers from CSX to Alphabet told investors in recent days that they plan to add people, a reversal after eighteen months of treating hiring as a last resort. The CEO of the HR platform Lattice said many companies stopped hiring junior staff on the assumption that AI agents would cover the work, then realized humans are still needed to work alongside the tools. Her line: having coding agents does not mean you stop hiring engineers, and AI sales agents still need salespeople.

    Separately, initial jobless claims fell to 187,000 for the week ending July 18, the lowest level since September 1969. I checked that against the Labor Department figures, and it holds across Bloomberg, Reuters, and CNN. The year opened with predictions of an AI jobs apocalypse. It is currently producing the fewest unemployment filings in nearly sixty years.

    Why this fits my argument

    What the reporting confirms is the narrow prediction. Cutting on anticipation, ahead of what the technology could actually deliver, was premature, and some of it is now being unwound. The rehiring is the anticipation bet reversing. When the country’s largest employers cut on the theory that agents would absorb the work, then hire back because the agents did not, that is a story about companies acting on a capability that was not there yet.

    What the reporting does not confirm is my deeper claim. My argument was that the real constraint is coordination, the cost of making capable systems work together and with the people around them. The Journal says companies are hiring because of cost, limitation, and uncertainty. It does not say they are hiring because their AI deployments fell apart at the seams. So take this as evidence that the premature-cut prediction was right, and as an open question on the coordination claim, which is the one still worth watching.

    Two caveats on the numbers. Economists describe this as a low-hire, low-fire market: layoffs are low, but hiring is soft too, and June’s dip in the unemployment rate owed partly to a shrinking workforce rather than a boom. And a single week of claims data is noisy, prone to summer seasonal swings.

    The line that spoke the truth

    The MIT labor economist Paul Osterman gave the Journal the most honest sentence in the piece. Asked whether companies need more people or fewer, he said no one has any idea. That uncertainty is the actual state of things.

    The jobs question was never how many humans the machines replace. It was whether an organization understands its own work well enough to know what it can safely hand off. Several did not, so they guessed, and some are now hiring back the people they let go. That is a coherence story, and it is the one I will keep following.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • The Subtitle I Almost Got Wrong

    Newsletter – Edition 2

    Those of you following my book’s journey from early on may have noticed that the subtitle has changed, and the reason is a small story about a mistake I almost made.

    A few weeks ago my publisher told me, reasonably, that a reader likely couldn’t tell what the book was about at a glance. He wanted “agentic AI” in the subtitle, both for clarity and for search. At the time it read “What Wins When Everyone Can Go Fast,” which I liked and which never once said AI.

    My instinct was to push back, and my first reason was not a good one. I didn’t want the book filed under yet another agentic-AI title when it’s a management book about the enterprise consequences of AI. That instinct had already cost me. Trying not to sound like an AI book, I’d written a subtitle that didn’t clearly signal what it was about.

    The reason I actually cared about came out when I started drafting replacements. I didn’t want to compete for the noisiest keyword in the market. There are thousands of things shouting “agentic AI” right now, and being one more voice in that crowd is not the same as being found. I wanted words that would still hold up after AI stops being novel.

    A quick test helped me sort the candidates. Say the subtitle out loud with “electricity” in place of “AI.” “The Last Advantage When Every Company Runs on Electricity” sounds a century old, which told me the phrasing was tied too tightly to the moment.

    I settled on “Coherence: The Competitive Advantage AI Can’t Buy.” It names AI, which was the fair part of my publisher’s request. It makes an argument instead of chasing a search term, which was the part I wasn’t willing to give up. The advantage AI can’t buy is the one your competitor can’t buy either, because you’re both shopping from the same shelf. What isn’t on that shelf is the whole point.

    The week in ideas

    Three posts from the past week.

    Good AI Governance Is Not the Same as Coherence. Australia’s directors just got one of the best AI governance guides I’ve read, and I spent the post explaining why the best version of the mainstream answer still misses the failure that will catch these boards. A governance apparatus works by review, and it reviews what reaches it. The incoherence that builds up between separately approved systems never comes up for a vote. And the human oversight everyone prescribes can pass its own audit while quietly going hollow, as a stretched review team waves through more and catches less. Real news from last week makes the point. Anthropic, the company that sells agentic AI, published a sober four-question checklist for deploying it safely: what untrusted content does the agent ingest, what can it do, what’s the blast radius, can you see what it’s doing. Four good questions. Every one of them inspects a single agent, one at a time. None of them can see the incoherence that accumulates in the space between agents that each passed. [Weigh in on LinkedIn…]

    The Moat Is Coherence. Kirkland & Ellis, the highest-grossing law firm in the world, is spending around half a billion dollars to build its own AI rather than rent what its rivals can rent. Its chairman put the logic in a line: widely available tools raise the floor for everyone, and the firm doesn’t get hired for the floor. The post works out what Kirkland is actually buying, which isn’t the model and isn’t the data, but the coherence that turns both into judgment a competitor can’t copy. The giveaway is the exclusivity clause. If the value were the technology, keeping it exclusive wouldn’t matter, because the technology is for sale to everyone anyway. [Weigh in on LinkedIn…]

    The Machine Proved It. Did It Do Mathematics? A digression, and my favorite of the three. An AI model recently disproved a conjecture the mathematician Paul Erdős posed in 1946, reaching across the field into tools no human specialist would have thought to try. The machine produced the proof. It did not choose the question. The post works through why understanding isn’t decoration but compression, the way a bounded human mind fits something enormous into the space of a single brain, and why the usual “trust the result you can’t follow” analogy from medicine breaks down in mathematics, which has no second way to verify a claim besides the proof itself. The closing point holds up under the whole argument. Nobody has built a system that decides which question is worth eighty years of human attention. [Weigh in on LinkedIn…]

    There’s one thread through all three. The machine can produce the output. A human still owns the part with no dashboard: choosing the question, holding the separate pieces together, judging whether the answer is any good. Kirkland is paying half a billion dollars to own that part. The governance guides keep prescribing a version of it that passes the audit and can’t see. The mathematicians are the ones saying out loud that they don’t yet know how to measure it. It’s the same thing the new subtitle names. The advantage isn’t the AI. It’s the coherence around it, and that isn’t for sale.

    Before you go

    One more item, because it belongs to the same story. SAP just closed its acquisition of Prior Labs and committed more than a billion euros to a lab that builds foundation models for structured data instead of text. The bet is that the untapped value in enterprise AI sits in the tables and databases a business actually runs on, not in another chatbot. I think that bet is right and incomplete in a familiar way. Point a powerful model at a fragmented, contradictory data estate and you get fast, confident reasoning over an incoherent picture, which is worse than slow reasoning, because the confidence hides the mess. Making the data coherent enough to trust is the half nobody can sell you.

    If the book is why you’re here, it’s Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. Those stories are where a good share of my ideas come from.

  • Safest Car on the Road, Yet Parks in the Fire Lane

    I read an article in the Wall Street Journal today about my hometown of Austin, and it made me laugh before it made me think. Since Waymo’s robotaxis arrived here in 2024, they have apparently collected $9,325 in parking tickets. Tow-away zones. Metered spots they never paid. A disabled space outside an elementary school. One that idled in front of a church garage for five minutes during Sunday service while parishioners waited. A resident’s complaint in the records reads, plainly, “there needs to be some way to get them to move.”

    These figures come from documents the Journal obtained through an open-records request, so I am relaying its reporting rather than confirming the numbers myself.

    As a number, $9,325 is nothing. Austin collected $6.3 million in parking fines in 2025 alone, so Waymo’s two-year total is a rounding error. Spread 83 citations across more than 300 cars over two years and the per-vehicle rate is low, probably lower than what a human-driven taxi fleet of the same size would rack up in the same window.

    While the dollar figure is trivial, the behavior behind it is not. It is also a near-perfect illustration of the argument I have been making.

    The obvious reading, and why it misses

    The easy version of this story is that the self-driving car is not ready. Look, it cannot even park. That reading is wrong, and the same article contains the reason.

    An independent analysis by the Insurance Institute for Highway Safety found that over more than 50 million driverless miles, Waymo’s crash involvement rate was 68 percent lower than that of human drivers. The hard problem, the one with lives attached, Waymo has solved to a level that beats us. Parking is where it stumbles.

    A system can be superhuman at its central task and fail at something a sixteen-year-old handles on the first day with a learner’s permit. This is jaggedness. Andrej Karpathy coined the term for the strange fact that a state-of-the-art model can solve a hard problem and then miss a trivial one, and a field experiment with 758 BCG consultants showed the same thing in the workplace: performance was excellent on tasks inside the model’s zone and worse on adjacent tasks that looked just as easy. The boundary is uneven, and it does not follow the difficulty ranking a human would draw. Driving safely across 50 million miles is the hard task the machine has mastered. Parking lawfully is the easy adjacent task it has not, and no amount of skill at the first predicts skill at the second. I have written about this shape once already this month. An AI office agent filled out seventeen forms in five minutes and then could not upload a file. Same jaggedness, different machine.

    The dangerous part is that the failure is invisible from the outside. Watch a car drive flawlessly for an hour and you will assume it can handle a parking lot, because that inference holds for humans. It does not hold here, and the assumption is where the trouble starts.

    The parking failure itself is not my thesis. Waymo will patch handicap-spot detection, and that particular fine will stop appearing. But notice what does not happen. Jaggedness does not get fixed. It moves. Patch the parking lot and the uneven edge shows up somewhere else nobody thought to check, because the unevenness comes from how the system learns, not from a single defect waiting to be found.

    That is one problem, and it is real. There is a second one in the same article, and it is not a version of the first. It is a different failure entirely, and it is the one my book is actually about.

    The failure that no model fixes

    Read the part of the article that is not about parking.

    On July 8, the National Highway Traffic Safety Administration sent autonomous-vehicle developers a letter demanding that their cars better follow instructions from first responders. The regulator’s complaint was that robotaxis often fail to recognize where they can stop without getting in the way. When a Waymo blocked an active railroad track in January 2025, an officer reported he had no choice but to have it towed. There was no other way to move it.

    It does not go away with better driving. A firefighter at a scene, a police officer at a closure, a resident at a blocked garage: each of them has authority over the situation but no means to direct the machine sitting in it. The human is formally in charge and practically helpless. I keep making one distinction in my book, and this is it in the physical world. Having authority over a system is not the same as having the capacity to intervene in it. The officer had every right to move that car. He had no lever to do it, so he called a tow truck.

    That gap is the coherence problem, and no amount of driving skill closes it. It is a problem of the connection between a capable system and the people who are supposed to be able to redirect it. You can make the car a better driver every quarter and leave that gap exactly where it is.

    Three hundred locally rational decisions

    Here is the detail in the article that matters most.

    Waymo runs more than 300 robotaxis in Austin. Between trips, the article says, the cars park themselves on public streets to stay near riders and avoid adding traffic. Each of those choices is sensible. Idling near likely demand cuts empty miles and shortens the next pickup. For the fleet, it is the right call every time.

    Now add up 300 right calls. You get 300 vehicles independently claiming curb space across one city, each optimizing for the fleet, none of them accountable for what they cost the curb in aggregate. The church-garage blockage was not one rude car. It was the predictable output of a fleet doing exactly what it was designed to do, measured against a shared resource that no one in the system is responsible for.

    This is the pattern I spend the book on. W. Edwards Deming showed it in factories long before any of this. Optimize each part on its own and the whole can still degrade, because the parts interact in ways no single part can see. A support agent and a billing agent inside a company can each be flawless and still act on contradictory assumptions about the same customer. Three hundred robotaxis can each park perfectly rationally and still congest a city. The mechanism is identical. The only new thing is that it now runs at the speed and scale of software, on a public street.

    Where the analogy breaks, and why the break is the interesting part

    My book is mostly about a different situation. Many systems, built by many teams, with no shared owner, colliding inside one company. Waymo is close to the opposite. One company, one software stack, one central fleet manager. The cars are not incoherent with each other. They are all perfectly coherent with Waymo’s goal. The incoherence is between the fleet and the city.

    That difference does not weaken the parallel. It sharpens it. Inside a single enterprise, the cost of local optimization eventually lands back on the enterprise itself. It pays its own complexity debt, later and with interest. In the robotaxi case, the company captures the efficiency and the public absorbs the cost. The fleet gets the shorter pickup times. The churchgoers get the blocked garage. The externality lands outside the firm, which means the firm has no natural reason to see it or price it.

    Which is where the parking ticket returns, transformed. The tickets are not the failure in this story. The tickets are the city’s answer to it. Austin cannot rewrite Waymo’s software, so it does the only thing available to an outsider. It attaches a dollar figure to each incoherent act and bills it back. A tow-away citation is a coordinating institution forcing a cost back onto the party that created it, because that party will not absorb a cost it cannot see on its own dashboard.

    In the book I call that pricing the incoherence, making the party that creates a coordination burden bear the cost it imposes on everyone else. It is one of only a few places you can intervene when you cannot redesign the system directly. Austin is doing it with a parking-enforcement officer and a public complaints database. It may be crude, but it is the right instinct. When you cannot fix the system, you can at least make it pay for the mess, so the incentive to stop finally reaches someone who can.

    The instinct is widely shared, which is its own small piece of evidence. I read the comments under the article, and the readers who were not busy mocking it reached, unprompted, for exactly this lever. Charge a flat $5,000 fee for a driverless tow. Impound the car and make the company pay storage, same as a person would. Hold them to the standard you and I are held to. Nobody in that thread proposed debugging Waymo’s curb detection, because none of them can. They proposed raising the price of the behavior, which is the one move available to an outsider who cannot see inside the system and cannot change it. Pricing is what is left when coordination is out of reach.

    A note on fairness, since it matters. Waymo pays these tickets like any other driver, and its spokesman said the company expects no special treatment. It contests some citations and has had a couple dismissed. None of that is evasion. It is a company behaving reasonably inside a system that has not yet given it a better way to behave. The point is not that Waymo is careless. The point is that even a careful, centrally managed, genuinely safer-than-human fleet produces coordination costs its own metrics will never show. That is the part that should worry anyone deploying autonomous systems anywhere.

    The version of this that has not happened yet

    One last thought, and I will flag it clearly as speculation rather than something the article reports.

    Today Austin has one large fleet parking itself on the curb. The same article names two more operators already here, Tesla’s Robotaxi and Amazon’s Zoox. Imagine the near future where three or four fleets, each centrally coherent, each optimizing its own vehicles against the same finite curb, all share one city. None of them is incoherent on its own terms. Each is a model citizen by its own dashboard. Together they compete for the same few feet of pavement outside the same church at the same 10 a.m. service, and no one owns the result.

    That is the multi-owner version of the trap, and it is the one that looks most like the enterprise problem I actually write about. Many capable systems, no shared view of the whole, a commons that quietly degrades while every participant is behaving well. When it arrives, the city will reach for the same tool it is using now, only harder. It will try to price the congestion, because pricing is what is left when you cannot coordinate the systems directly and you cannot see inside any of them.

    There is a sharper edge to a single fleet that is worth one more sentence, because it cuts against the intuition that central control is safer. A fleet of 300 identical cars does not only optimize together. It fails together. Every vehicle runs the same software and leans on the same positioning inputs, so a single upstream fault does not hit one car, it hits all of them at once and in the same way. I made this point about software agents in a recent post: two agents drawn from the same model share the same blind spot, so the redundancy between them is nominal. Here it is 300 machines sharing one blind spot instead of two. Homogeneity buys clean coordination on a good day and correlated failure on a bad one. That is not an argument against central control. It is a reminder that the thing which makes a fleet coherent is the same thing that can make it fail in unison.

    The safest car on the road parks in the fire lane. The fleet that adds no traffic blocks the garage. Every decision was locally correct, and the street got worse anyway. That is not a story about cars. It is the story of what happens to any organization, or any city, that fills up with capable systems faster than it builds the means to keep them coherent.

    The gap between capable systems and coherent ones is the subject of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.