Author: Madhu Shashanka

  • Could a Machine Have Had Darwin’s Idea?

    In November 2020 I started a thread on X. I have barely used the platform these past couple of years, and I only came back to this thread recently, by accident, when some of my own old posts surfaced in front of me. Reading them in one sitting was strange, because a thread I had added to piecemeal over five years turned out to have been about one question the whole time.

    It opened with Darwin. Someone I followed had posted, amazed, about what evolution manages to build, and I wrote that what amazed me was something else: the intellectual leap required to infer the theory in the first place, from nothing but empirical observation and years of focused study. I said I was not sure I fully appreciated Darwin’s intellect and dedication. Then, in the next post, I said the thing the whole thread has been chasing ever since. That kind of leap, I wrote, is the kind of intelligence AI should be aiming for, not recognizing cat faces or copying tasks people already do well.

    That was the bar I set in 2020, and it was a high one. Darwin spent decades buried in finches and barnacles and pigeon breeders and fossil beds, a mountain of unconnected observation, and then made a leap: one idea, natural selection, that reorganized all of it at once. That inductive move, from a heap of messy particulars to the principle that explains them, struck me then as the most formidable act an intellect can perform, and the part of real science I was least sure a machine could touch. So I kept a list, adding to it whenever I came across a case of AI looking like it was helping advance science rather than just crunch it, to watch whether anything ever cleared the bar.

    Reading the thread back now, it traces an arc I did not plan, and it worried at the right problem from the start. Within days, in November 2020, I posted a New York Times piece on Max Tegmark’s group, whose neural network had recovered a hundred physics equations from raw data, and I pulled out the catch that the physicists themselves named. Tegmark was candid that the machine could retrieve the formulas but not yet the deep principles beneath them, the quantum uncertainty or the relativity that would explain why the formula holds. And Jesse Thaler, the MIT physicist directing the new AI-and-physics institute, put his finger on why. AI wins at games because a game has a well-defined notion of success. “If we could define what success means for physical laws,” he said, “that would be an incredible breakthrough.” Proposing the theory, in other words, was the hard part, and it was hard precisely because you could not say in advance what would count as getting it right.

    The entries that moved me most in those early days were not about AI at all. They were about humans doing the thing I wanted to see a machine do. Around the same time I was reading about Tibor Gánti, the Hungarian biologist who deduced from first principles what the simplest possible living thing must be: a metabolism, a way to store information, and a membrane, three systems that have to be coupled or the organism dies. I wrote at the time that this was the kind of inductive thinking we should be teaching children. Later I followed a related idea, assembly theory, which proposes a single elegant handle on complexity: count the minimum number of steps needed to build a molecule, and use that number to test for the presence of life on other worlds. A theory reduced to one measurable quantity, aimed at one of the hardest questions there is. It may not survive contact with the evidence; I later noted that a study found some minerals scoring above the threshold the theory sets for life. But right or wrong, it was the move I admired, the leap from observation to a principle sharp enough to be tested.

    Then the machine examples accumulated. In December 2021, a Nature paper where neural networks guided the intuition of mathematicians toward new conjectures in knot theory, the machine surfacing the pattern and the human still making the leap. In March 2022, a system that rediscovered Newton’s law of gravitation from the motion of the planets and wrote it back out as a symbolic equation. Then more symbolic regression, pulling laws from data. Then, in February 2025, the entries that made me sit up: Google’s AI co-scientist, generating and ranking novel research hypotheses on its own, and Evo-2, a model that does not just read genomes but writes them. And I drew a line I did not fully understand at the time. I noted that the most exciting work was different from the systems that merely “generate plausible hypotheses from an extremely large space of possibilities.”

    Five years of watching, and that distinction turned out to be the whole story.

    The step that was supposed to have no method

    For a century, the standard account of science drew a hard line between two acts.

    One is coming up with the idea. The other is checking whether the idea is true. Karl Popper named these the context of discovery and the context of justification, and he was blunt about which one belonged to philosophy. Testing a hypothesis has a logic. Having one does not. In his words, the act of conceiving a theory “neither calls for logical analysis nor is susceptible of it.” The initial leap from a pile of observations to this might be why was, he thought, a matter for psychology, not method. A hunch. The unteachable part.

    That is where the whole romance of science lived, and Darwin’s leap is its patron saint. Kekulé dreaming the benzene ring as a snake biting its tail. Fleming noticing the one culture plate that had gone wrong in an interesting way. Darwin holding twenty years of specimens in his head until they resolved into a single idea. We told these stories because the leap seemed to come from nowhere, and coming from nowhere was the point. You could train someone to run an experiment. You could not train the hunch.

    The systems in my thread are automating the hunch.

    Not perfectly, and not everywhere. But the co-scientist does not summarize the literature and hand you a reading list. It proposes mechanisms no one has written down, argues them against itself, and ranks the survivors. Evo-2 does not retrieve a gene; it composes one. Whatever you want to call that, it is happening on the discovery side of Popper’s line, in the territory he declared off-limits to method. The unteachable step is being done by a machine that was, in fact, taught.

    What actually got cheap

    Here is where my thread stops being a highlight reel and starts being an argument, because the interesting question is not whether machines can generate hypotheses. They plainly can. The question is what that does to the rest of science.

    The mathematician Noah Giansiracusa has a name for the pattern, which I take up at more length in my book: carpet bombing. When generation gets cheap, you stop being clever about producing candidates and start producing all of them, then sort. It is how AI does mathematics, throwing enormous numbers of attempts at a problem where checking each one is fast. Hypothesis generation is now carpet bombing pointed at nature. The co-scientist can produce more plausible, well-argued, literature-grounded hypotheses in an afternoon than a lab could dream up in a year.

    And that is exactly where the trouble starts, because a hypothesis is not a proof. It cannot be checked in an afternoon. It has to be checked against the world, and the world runs on its own clock.

    When generation was expensive, the scarce, precious act was having the good idea, and verification, while never easy, was not the binding constraint. A scientist had three hypotheses worth testing and a career to test them in. Reverse that. Now the machine hands you three hundred plausible hypotheses, and the binding constraint is the wet lab, the clinical trial, the telescope time, the years. Generation raced ahead. Verification did not move at all, because verification in science is not a faster model. It is reality, taking as long as reality takes.

    This is the same shape I keep finding everywhere AI touches real work, and I wrote about its purest form in mathematics in an earlier piece. Cheap generation does not remove the bottleneck. It moves it downstream and makes it the whole game.

    The field is already learning this the hard way

    You do not have to take the argument on faith, because the correction is already arriving, and it is arriving in the most useful form: from the people who built the tools.

    Google’s co-scientist reached Nature in 2026, with real wet-lab validation in a handful of biomedical cases. Impressive, and I do not want to wave it away. But when an independent researcher carefully re-implemented the system and ran it hard, the finding was sobering. The pipeline reliably produces hypotheses. Whether it actually improves on the underlying model’s raw guesses was not reproducible from one run to the next, and across dozens of attempts on one disease, not a single one of its generated hypotheses matched the paper’s own headline discoveries. The machine is a fountain of plausible ideas. Plausible is not the same as true, and telling them apart is still the expensive part.

    There is a sharper cautionary tale two years older. In 2023 Google reported that around forty new materials had been discovered and synthesized with the help of one of its AI systems. It was held up as a landmark. Then outside chemists went through the results, and an independent analysis concluded that not one of them was actually a net-new material. The generator worked. The verification, done properly and after the fact by humans, is what separated the discovery from the illusion of one. Every honest account of these systems now carries the same caveat, in the developers’ own words: careful experimental validation, peer review, and independent scrutiny are what turn a generated candidate into knowledge.

    That caveat is not a footnote. It is the job.

    Which part was ever the science

    So let me go back to my thread, and to the line I drew without fully appreciating it.

    I think I was reaching for this, and the clue was in Thaler’s line about defining success. The systems I found most exciting were not the ones that produced the most hypotheses. They were the ones tied to a way of checking, the model that rediscovered gravity and could be tested against known physics, the mathematical work where a conjecture could be pursued to a proof. The ones that unsettled me were the pure generators, magnificent at producing possibilities and silent on which ones were real. The 2020 worry and the 2025 worry are the same worry. A physical law you cannot define success for, and a hypothesis you cannot yet verify, are the same problem.

    Popper drew his line to protect justification. He wanted to say that the logic of science lived in the testing, and that the having-of-ideas, however romantic, was not where rigor lived. The machines have now inverted his world in the most ironic way possible. They have automated the part he thought had no method, the hunch, and in doing so they have made the part he cared about, the checking, more valuable than it has ever been. When hunches were scarce, verification could feel like bookkeeping. Now that hunches are infinite and nearly free, verification is the only thing standing between a lab and a year spent chasing a beautifully argued hypothesis that was never going to be true.

    There is a clue to this in something Andrew Wiles once said, that it is bad to have too good a memory if you want to be a mathematician. It sounds backward until you see what he means. The mathematical gift was never recall. It was compression, the knack for throwing away almost everything and keeping the one idea that organizes the rest, which is roughly how Jürgen Schmidhuber defines insight: a better, shorter way to predict what you have seen. A machine with perfect memory and unlimited generation has exactly the strength Wiles warns against and not yet the one he prizes. It can hold everything. It cannot yet tell what to forget.

    Which returns me to the question I started the thread to answer. Could a machine ever do what Darwin did?

    I think the honest answer is now a qualified yes, and it is qualified in a way I did not expect. A system can hold a mountain of observation and propose the organizing idea. It can make the inductive leap I was so sure was ours alone. But watching it happen, I realized I had misjudged where Darwin’s genius actually sat. The leap to natural selection was extraordinary, but the leap was not the science. The science was in the twenty years, in Darwin knowing which single idea out of the dozens he entertained was worth a life’s defense, and in his relentless testing of it against every objection he could invent. The hunch was the cheaper half but we just could not see that while hunches were rare.

    So the machines got the hunches, and they are welcome to them. What Darwin had that they still do not is the judgment to know which hunch was worth everything, and the patience to spend years finding out if he was wrong. That part did not get automated. It got scarcer, and more valuable, than it was when the ideas were hard to come by.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Can Frontier AI Outdo MBAs?

    Three of the top business schools in the country just tested frontier AI on the analytical work their MBAs are trained to do, and the models scored in the high eighties. If you run a company, that number is coming for you soon, probably in a deck that recommends cutting a layer of analysts. So it is worth being exact about what it measures, because the exact answer is more useful, and more limited, than the headline.

    The paper is BusinessCaseBench, from researchers at Wharton, Carnegie Mellon, and Harvard Business School. They drew 615 questions from real business school cases across eighteen disciplines, from strategy and finance to leadership and ethics. Each model read a case cold and wrote its analysis. A separate grader then compared that analysis against the reference solution the instructor had written, item by item. The models never saw the reference. As far as the model was concerned, it faced an open business question with no answer attached, and it answered well. Under the main metric, Claude Sonnet 4.6 covered about 88 percent of what the instructor’s solution contained, and GPT-5.4 about 87.

    The model was not helped by the answer key. It could not see the rubric, did not use it, and produced its analysis from the case alone, the way a consultant works from a brief. These were, from the model’s side, genuinely open questions. So it is fair to expect that a model which writes strong analyses on 615 unseen cases will write a strong analysis on the 616th, which is your real one.

    The problem is that “the model will write a similar analysis” and “you will get a similar result” are different claims, and only the first one is what the benchmark tested.

    A score is a comparison

    “The model scored 88 percent” is not a fact about the model’s answer by itself. It is a fact about that answer measured against a standard. Two things had to exist for the number to exist: the analysis the model wrote, and the instructor’s solution it was checked against. The score is the relationship between them.

    Now move that setup to your company. The model can still write the analysis. But the standard it gets checked against does not exist. The 88 was a statement about how well the answer matched a known-good answer.

    This is not word games. It is the difference between “the model is competent” and “you can trust the output.” The first is about the answer. The second is about checking the answer, and checking requires a standard. The benchmark supplied the standard. Your hardest decisions do not.

    Knowable-but-hidden is not the same as unknowable

    The tempting reply is that the real world is just the benchmark with the answer hidden. The model handled hidden answers fine; a real decision is one more hidden answer.

    But the benchmark’s answers were not hidden. They were knowable in the first place. The professor had already worked the case. A correct answer existed; the model simply was not shown it. That is a different situation from the one you are in when you decide whether to enter a market or restructure a division. There, no correct answer exists yet. It has not been written by anyone, because the outcome that would settle it is years away, the criteria for “good” are contested by the people in the room, and there is no counterfactual to check the decision against even after the fact.

    A graded case has a knowable answer the model didn’t see. A live strategic decision has an unknowable answer that does not exist to be seen. A student who scores 88 on a past exam she took blind will likely score about 88 on the next past exam. But it does not follow that she will make good venture bets, even though both feel like hard open-ended judgment, because a venture bet has no marking scheme, then or later. The model is the student. The benchmark is the past exam. Your boardroom is the venture bet.

    So the benchmark is strong evidence for a real claim: frontier models are good at producing structured business analysis, and getting better fast. It is not evidence for the claim that the score predicts a good outcome on decisions whose standard has to be invented rather than looked up. Inventing that standard, deciding what a good answer to your actual question would even need to contain, is the judgment.

    That argument stands even if the models are excellent.

    The model is fully right about half the time

    The researchers scored the answers two ways. The headline 88 percent is partial credit: how much of the instructor’s checklist each answer covered. Then they ran a stricter count. On how many questions did the answer satisfy the entire checklist, every item, no gaps? That fell to roughly half.

    On cases where a correct answer was knowable, the leading model produced a complete answer about one time in two. The authors put the point in their own title for that result: these are drafts, not verdicts.

    Now extrapolate honestly. If you carry this model into the wild and expect “similar performance,” you are also carrying the incompleteness. The typical output is strong and missing something at the same time. In the study, a human holding the instructor’s solution caught the missing half. In your firm, if you removed the person who could catch it, the missing half is still missing and nothing catches it.

    Grader didn’t check for what shouldn’t be there

    The grader verifies whether each expected point is present. By construction, it does not scan the answer for confident, invented, or wrong material that sits outside the checklist. A response can hit the expected points and also assert three plausible fabrications and still score well, because nothing in the method is looking for the fabrications.

    In the study that blind spot is harmless, because a grader with the solution ignores the extra material. In your company it is the whole risk. The fabricated line rides along inside a well-organized analysis, unflagged, and the reader who could catch it is the domain expert the strong score seemed to make optional.

    Why “drafts, not verdicts” is the expensive finding

    A draft that is 88 percent right sounds like a bargain, and sometimes it is. But it moves the work rather than removing it. Obvious garbage you would catch. Verified truth you could trust. The strong, incomplete, possibly-embellished draft is the expensive case, because it earns your trust on the parts you can see and hides its gaps in the part you would have to already know the answer to find. Catching what it left out requires someone who knows what a complete answer contains, which is exactly the expertise the draft appeared to retire.

    So the generation got cheap and the checking did not, and on this kind of work the checking cannot be sampled down. You cannot review only the flawed answers, because the flawed ones are the ones that look fine. You review all of them, or you review none and call it oversight. An organization that reads the 88 as license to remove the reviewers has not automated the analysis. It has removed its own ability to tell when the analysis is wrong.

    What it means if you run something

    Read correctly, the benchmark is good news. Frontier models are genuinely strong at producing structured business analysis, and that is real leverage you should use. The error is reading a score-against-a-known-standard as a readiness-to-deploy-without-a-standard.

    Two questions to ask of any AI you are about to trust with judgment work.

    Does this task come with a knowable answer, or is defining the answer the actual job? Where a defensible right answer exists, straightforward analysis with a standard you could write down, let the model run, and expect it to perform in the wild about as it did on the bench. That extrapolation is fair, and worth taking. Where the standard itself has to be invented and argued, the model is drafting, and a person still owns the decision.

    And who checks the drafts, and do they know enough to catch what a confident draft leaves out? If the answer is nobody, or nobody who could tell, the high score is measuring something you will not actually get.

    The models are good at the framed question. Your hardest problems arrive unframed. The competitive advantage was never in answering the case. It was in knowing which case you were actually in.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Sensing Is More Than Measurement

    The Financial Times reported this week on an internal Amazon presentation held July 28. Engineers walked staff through a set of AI cost overruns. The largest used Anthropic’s Claude Sonnet to match author details against product listings on Amazon’s retail site. It ran 860 percent over budget, cost $1.8 million, and never shipped. Engineers said mistakes that were once trivially cheap had become “catastrophically expensive”.

    A financial auditing tool ran about $541,000 over. A logistics project meant to reduce delivery times ran about $134,000 over. Roughly $2.5 million across the three.

    News coverage led with the money. Several outlets called it a coding task. Matching author details to listings is data reconciliation, the kind of work nobody watches.

    The project ran for five months before anyone caught it.

    The data was there

    A senior Amazon employee told the FT that it is difficult to figure out how much anything AI-related costs. Said by someone at the company that runs the cloud everyone else buys AI on.

    Token spend is metered, priced publicly, and billed monthly. Every token that project consumed appeared on an invoice. Engineers attributed the overruns partly to the shift from flat subscriptions to token-based billing, where costs climb whenever a task generates more activity than expected.

    Few things inside a large enterprise are more thoroughly instrumented than a cloud bill. Amazon had the numbers for five months and stayed unaware of them.

    Sensing takes more than measurement. Something has to compare the number against an expectation, notice the gap, and route it to someone who can act while acting is still cheap. Amazon had the number. The comparison and the route were missing.

    No dashboard would have closed this. Someone had to decide that aggregate token spend against declared intent is a thing the company watches, and then own the watching.

    March incidents

    Amazon’s retail website took four high-severity incidents in a single week in early March, including a six-hour failure that locked customers out of checkout, account information, and pricing.

    An internal document prepared for the review meeting identified GenAI-assisted changes as a factor in a pattern of incidents going back to Q3. That reference was deleted before the meeting, according to the FT, which saw both versions. Amazon disputed the reporting and said only one incident involved AI directly, with the root cause an engineer acting on inaccurate advice an AI agent had inferred from an outdated internal wiki.

    Amazon’s response was a 90-day code safety reset across 335 critical retail systems and mandatory senior-engineer sign-off on AI-assisted code from junior and mid-level engineers.

    Work backward from July. Five months of undetected spending starts around February or March. I cannot confirm the detection date, so treat the overlap as inference. Even without it, the shape holds. Amazon added review gates on AI-assisted changes to critical systems while a cost failure accumulated invisibly on a job nobody would call critical.

    Constraint depends on detection. You cannot cap, contain, or price what you cannot see. Amazon reached for the second without the first, which happens because approval steps are visible to leadership and instrumentation is not.

    KiroRank

    Amazon ran an internal leaderboard called KiroRank that ranked employees by AI usage. Staff responded with what they called tokenmaxxing, deliberately inflating consumption to climb the rankings. Amazon discontinued it.

    Every deployment a team builds imposes cost on everyone else. Another surface to watch, another dependency to reconcile, another system someone who did not build it has to understand. The team keeps the benefit while the organization bears the cost. The remedy is to price that burden back to the team creating it.

    Amazon built a price signal pointing the wrong way. Ranking people by consumption pays them to consume. An internal metric carrying status and no cost gets gamed, and this one did.

    The presentation’s own recommendations now include avoiding leaderboards that reward token consumption, and checking whether higher token usage produces useful output.

    The objection

    Amazon frames these as isolated examples of teams learning from one another, and says cherry-picking them does not reflect how teams across the company use AI. For a company its size, a seven-figure surprise is a rounding error.

    The reporting also lacks a base rate. Nobody has said how many AI projects came in on budget for every one that blew up.

    Five months of invisibility still belongs to the control architecture rather than the budget. The same architecture at a company with a $12 million annual AI budget produces the same five months and a different outcome. Amazon can absorb what it cannot see.

    What it costs to fix

    Spending on sensing is bounded and knowable in advance. You can price the instrumentation, the ownership, and the review cadence before committing. The incoherence it prevents accrues silently and surfaces only once addressing it is no longer optional.

    Amazon paid roughly $2.5 million across three disclosed projects. Other companies will meet the same failure without the revenue to absorb it.

    The argument here about sensing, constraint, and priced externalities runs through my book, Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Why Ford Rehired

    Ford has added more than 350 experienced engineers over the past three years after its automated quality systems failed to deliver the results the company expected. Inside Ford they are called gray beards. Some are former Ford employees. Others came from suppliers.

    Ford says it added the specialists “through internal promotions or new talent” to work alongside newer team members. Headlines have called it rehiring people AI replaced but Ford has not used that word, and the reporting does not establish that these specific roles were cut.

    Charles Poon, Ford’s vice president of vehicle hardware engineering, told reporters that AI is a fantastic tool and only as good as the information used to train it. He was more direct about the error: “Mistakenly, we thought that by just introducing artificial intelligence and ingesting the design requirements that we had, that would produce a high-quality product.”

    Poon also said Ford had not paid enough attention in prior years to the experience of its most knowledgeable engineers, the ones who had been through many product cycles.

    Kumar Galhotra, Ford’s chief operating officer, said the company had been relying more and more on automated quality systems before recognizing the approach was not working.

    Ford then ranked first among mainstream brands in the 2026 J.D. Power U.S. Initial Quality Study, its first time since 2010.

    Ford’s explanation

    Ford describes a training data problem. Poon’s version: enhancing the automation and machine learning tools required making sure they were trained by the most experienced individuals.

    Part of that holds up plainly. The specialists do reprogram the AI tools that fell short.

    The jobs

    A missing corpus has a fix with an end date. Sit the veterans down, extract the failure modes they carry, feed the models, thank them.

    But Ford built something with no end date. The specialists run mandatory meetings on quality concerns and hunt for failure points before a part reaches the plant floor. They conduct regular design reviews to identify problems before vehicles reach production. They train junior staff. Galhotra put the shift as moving from a find-and-fix mentality to preventing issues before they occur.

    A standing design review is a verification loop the organization has decided to keep running.

    Where the corpus story runs out

    Tacit knowledge can be captured up to a point. What a senior engineer knows about how a joint fails under a particular thermal cycle can be written down, and should be.

    Judgment applied to an unanticipated case cannot. A reviewer looks at a novel configuration and says it will not hold, for reasons that emerge from the thing in front of them. Enumerate those cases ahead of time and you would have automated the review already.

    The corpus framing implies a completion state, where enough capture makes the humans optional. Ford’s remedy points elsewhere. Mandatory and recurring is what you build once you have concluded the checking does not stop.

    Automating a quality inspection function means automating a verifier, which removes the capacity to tell whether the automation works.

    Ford’s specialists hold the ability to tell the machine it is wrong.

    The order

    Poon’s admission about prior years is the sharpest thing either executive said. Ford’s assumption was reasonable and its sequence was wrong. Capture the expertise, then automate, and the program is defensible. Automate on the assumption that design requirements are sufficient, and you spend three years buying judgment back from suppliers and internal promotions.

    Each step is the precondition for the next. Skipping one relocates its cost to a later point, larger, with fewer options available. Ford turned a knowledge problem into a three-year staffing program, and that program ran alongside more than $1 billion in expected warranty and material costs this year and a quality reputation to repair.

    Mentorship

    Ford is explicit that mentorship is part of the assignment. The specialists work alongside newer team members and train junior staff who never absorbed the institutional knowledge.

    This is how senior judgment gets built. Less experienced people work real problems while someone holding the judgment watches and corrects. The routine cases are the training ground, and automation takes them first.

    An organization that automates routine work and then loses its seniors breaks the pipeline at both ends. Nobody holds the judgment and nobody acquires it. Ford is paying to rebuild both.

    Headcount models calculate savings against the cost of the people. The judgment training pipeline appears nowhere in the model.

    Ford is still deploying AI

    On an autumn 2025 earnings call, Galhotra said Ford was systemically deploying AI across the entire industrial system, including 900 AI-powered cameras across its plants to detect quality issues at the source. Ford kept the cameras. Jim Farley told Bloomberg TV that Ford has AI tools for vision systems, and that most of it comes down to team members paying attention to small details.

    Ford topped the mainstream J.D. Power rankings with the automated systems still running. Experienced people now sit between the systems and the product.

    Experienced engineers sat on Ford’s books as execution capacity, a cost line. Their function was judgment, which is what makes execution capacity safe to deploy.

    Ford is not unusual in getting the order wrong. CNBC reported Robert Half data showing 32 percent of U.S. hiring managers eliminated a role primarily because of AI and later rehired for the same or a similar position. Robert Half’s own summary puts it as more than 3 in 10, and notes the two most common reasons given: the role required institutional knowledge or context AI could not replace, and it involved relationship management AI could not replicate.

    The arguments here about judgment as the layer that cannot be purchased, and about sequence in organizational automation, run through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Early Signs of Rehiring

    A short update to an earlier post on the AI jobs debate.

    A few weeks ago I argued that the AI jobs debate was asking the wrong question. My point was that the layoffs of the past year were mostly a bet, not a result. Companies were cutting in anticipation of what AI would let them do, ahead of the implementation that would justify the cut. Cut first, capture the savings later, and hope the two line up.

    The Wall Street Journal reported this week that, for a lot of large employers, they did not line up.

    What the article says

    Big companies are hiring again. The Journal reports that employers from CSX to Alphabet told investors in recent days that they plan to add people, a reversal after eighteen months of treating hiring as a last resort. The CEO of the HR platform Lattice said many companies stopped hiring junior staff on the assumption that AI agents would cover the work, then realized humans are still needed to work alongside the tools. Her line: having coding agents does not mean you stop hiring engineers, and AI sales agents still need salespeople.

    Separately, initial jobless claims fell to 187,000 for the week ending July 18, the lowest level since September 1969. I checked that against the Labor Department figures, and it holds across Bloomberg, Reuters, and CNN. The year opened with predictions of an AI jobs apocalypse. It is currently producing the fewest unemployment filings in nearly sixty years.

    Why this fits my argument

    What the reporting confirms is the narrow prediction. Cutting on anticipation, ahead of what the technology could actually deliver, was premature, and some of it is now being unwound. The rehiring is the anticipation bet reversing. When the country’s largest employers cut on the theory that agents would absorb the work, then hire back because the agents did not, that is a story about companies acting on a capability that was not there yet.

    What the reporting does not confirm is my deeper claim. My argument was that the real constraint is coordination, the cost of making capable systems work together and with the people around them. The Journal says companies are hiring because of cost, limitation, and uncertainty. It does not say they are hiring because their AI deployments fell apart at the seams. So take this as evidence that the premature-cut prediction was right, and as an open question on the coordination claim, which is the one still worth watching.

    Two caveats on the numbers. Economists describe this as a low-hire, low-fire market: layoffs are low, but hiring is soft too, and June’s dip in the unemployment rate owed partly to a shrinking workforce rather than a boom. And a single week of claims data is noisy, prone to summer seasonal swings.

    The line that spoke the truth

    The MIT labor economist Paul Osterman gave the Journal the most honest sentence in the piece. Asked whether companies need more people or fewer, he said no one has any idea. That uncertainty is the actual state of things.

    The jobs question was never how many humans the machines replace. It was whether an organization understands its own work well enough to know what it can safely hand off. Several did not, so they guessed, and some are now hiring back the people they let go. That is a coherence story, and it is the one I will keep following.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • The Subtitle I Almost Got Wrong

    Newsletter – Edition 2

    Those of you following my book’s journey from early on may have noticed that the subtitle has changed, and the reason is a small story about a mistake I almost made.

    A few weeks ago my publisher told me, reasonably, that a reader likely couldn’t tell what the book was about at a glance. He wanted “agentic AI” in the subtitle, both for clarity and for search. At the time it read “What Wins When Everyone Can Go Fast,” which I liked and which never once said AI.

    My instinct was to push back, and my first reason was not a good one. I didn’t want the book filed under yet another agentic-AI title when it’s a management book about the enterprise consequences of AI. That instinct had already cost me. Trying not to sound like an AI book, I’d written a subtitle that didn’t clearly signal what it was about.

    The reason I actually cared about came out when I started drafting replacements. I didn’t want to compete for the noisiest keyword in the market. There are thousands of things shouting “agentic AI” right now, and being one more voice in that crowd is not the same as being found. I wanted words that would still hold up after AI stops being novel.

    A quick test helped me sort the candidates. Say the subtitle out loud with “electricity” in place of “AI.” “The Last Advantage When Every Company Runs on Electricity” sounds a century old, which told me the phrasing was tied too tightly to the moment.

    I settled on “Coherence: The Competitive Advantage AI Can’t Buy.” It names AI, which was the fair part of my publisher’s request. It makes an argument instead of chasing a search term, which was the part I wasn’t willing to give up. The advantage AI can’t buy is the one your competitor can’t buy either, because you’re both shopping from the same shelf. What isn’t on that shelf is the whole point.

    The week in ideas

    Three posts from the past week.

    Good AI Governance Is Not the Same as Coherence. Australia’s directors just got one of the best AI governance guides I’ve read, and I spent the post explaining why the best version of the mainstream answer still misses the failure that will catch these boards. A governance apparatus works by review, and it reviews what reaches it. The incoherence that builds up between separately approved systems never comes up for a vote. And the human oversight everyone prescribes can pass its own audit while quietly going hollow, as a stretched review team waves through more and catches less. Real news from last week makes the point. Anthropic, the company that sells agentic AI, published a sober four-question checklist for deploying it safely: what untrusted content does the agent ingest, what can it do, what’s the blast radius, can you see what it’s doing. Four good questions. Every one of them inspects a single agent, one at a time. None of them can see the incoherence that accumulates in the space between agents that each passed. [Weigh in on LinkedIn…]

    The Moat Is Coherence. Kirkland & Ellis, the highest-grossing law firm in the world, is spending around half a billion dollars to build its own AI rather than rent what its rivals can rent. Its chairman put the logic in a line: widely available tools raise the floor for everyone, and the firm doesn’t get hired for the floor. The post works out what Kirkland is actually buying, which isn’t the model and isn’t the data, but the coherence that turns both into judgment a competitor can’t copy. The giveaway is the exclusivity clause. If the value were the technology, keeping it exclusive wouldn’t matter, because the technology is for sale to everyone anyway. [Weigh in on LinkedIn…]

    The Machine Proved It. Did It Do Mathematics? A digression, and my favorite of the three. An AI model recently disproved a conjecture the mathematician Paul Erdős posed in 1946, reaching across the field into tools no human specialist would have thought to try. The machine produced the proof. It did not choose the question. The post works through why understanding isn’t decoration but compression, the way a bounded human mind fits something enormous into the space of a single brain, and why the usual “trust the result you can’t follow” analogy from medicine breaks down in mathematics, which has no second way to verify a claim besides the proof itself. The closing point holds up under the whole argument. Nobody has built a system that decides which question is worth eighty years of human attention. [Weigh in on LinkedIn…]

    There’s one thread through all three. The machine can produce the output. A human still owns the part with no dashboard: choosing the question, holding the separate pieces together, judging whether the answer is any good. Kirkland is paying half a billion dollars to own that part. The governance guides keep prescribing a version of it that passes the audit and can’t see. The mathematicians are the ones saying out loud that they don’t yet know how to measure it. It’s the same thing the new subtitle names. The advantage isn’t the AI. It’s the coherence around it, and that isn’t for sale.

    Before you go

    One more item, because it belongs to the same story. SAP just closed its acquisition of Prior Labs and committed more than a billion euros to a lab that builds foundation models for structured data instead of text. The bet is that the untapped value in enterprise AI sits in the tables and databases a business actually runs on, not in another chatbot. I think that bet is right and incomplete in a familiar way. Point a powerful model at a fragmented, contradictory data estate and you get fast, confident reasoning over an incoherent picture, which is worse than slow reasoning, because the confidence hides the mess. Making the data coherent enough to trust is the half nobody can sell you.

    If the book is why you’re here, it’s Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone on the list gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me how it went. Those stories are where a good share of my ideas come from.

  • Safest Car on the Road, Yet Parks in the Fire Lane

    I read an article in the Wall Street Journal today about my hometown of Austin, and it made me laugh before it made me think. Since Waymo’s robotaxis arrived here in 2024, they have apparently collected $9,325 in parking tickets. Tow-away zones. Metered spots they never paid. A disabled space outside an elementary school. One that idled in front of a church garage for five minutes during Sunday service while parishioners waited. A resident’s complaint in the records reads, plainly, “there needs to be some way to get them to move.”

    These figures come from documents the Journal obtained through an open-records request, so I am relaying its reporting rather than confirming the numbers myself.

    As a number, $9,325 is nothing. Austin collected $6.3 million in parking fines in 2025 alone, so Waymo’s two-year total is a rounding error. Spread 83 citations across more than 300 cars over two years and the per-vehicle rate is low, probably lower than what a human-driven taxi fleet of the same size would rack up in the same window.

    While the dollar figure is trivial, the behavior behind it is not. It is also a near-perfect illustration of the argument I have been making.

    The obvious reading, and why it misses

    The easy version of this story is that the self-driving car is not ready. Look, it cannot even park. That reading is wrong, and the same article contains the reason.

    An independent analysis by the Insurance Institute for Highway Safety found that over more than 50 million driverless miles, Waymo’s crash involvement rate was 68 percent lower than that of human drivers. The hard problem, the one with lives attached, Waymo has solved to a level that beats us. Parking is where it stumbles.

    A system can be superhuman at its central task and fail at something a sixteen-year-old handles on the first day with a learner’s permit. This is jaggedness. Andrej Karpathy coined the term for the strange fact that a state-of-the-art model can solve a hard problem and then miss a trivial one, and a field experiment with 758 BCG consultants showed the same thing in the workplace: performance was excellent on tasks inside the model’s zone and worse on adjacent tasks that looked just as easy. The boundary is uneven, and it does not follow the difficulty ranking a human would draw. Driving safely across 50 million miles is the hard task the machine has mastered. Parking lawfully is the easy adjacent task it has not, and no amount of skill at the first predicts skill at the second. I have written about this shape once already this month. An AI office agent filled out seventeen forms in five minutes and then could not upload a file. Same jaggedness, different machine.

    The dangerous part is that the failure is invisible from the outside. Watch a car drive flawlessly for an hour and you will assume it can handle a parking lot, because that inference holds for humans. It does not hold here, and the assumption is where the trouble starts.

    The parking failure itself is not my thesis. Waymo will patch handicap-spot detection, and that particular fine will stop appearing. But notice what does not happen. Jaggedness does not get fixed. It moves. Patch the parking lot and the uneven edge shows up somewhere else nobody thought to check, because the unevenness comes from how the system learns, not from a single defect waiting to be found.

    That is one problem, and it is real. There is a second one in the same article, and it is not a version of the first. It is a different failure entirely, and it is the one my book is actually about.

    The failure that no model fixes

    Read the part of the article that is not about parking.

    On July 8, the National Highway Traffic Safety Administration sent autonomous-vehicle developers a letter demanding that their cars better follow instructions from first responders. The regulator’s complaint was that robotaxis often fail to recognize where they can stop without getting in the way. When a Waymo blocked an active railroad track in January 2025, an officer reported he had no choice but to have it towed. There was no other way to move it.

    It does not go away with better driving. A firefighter at a scene, a police officer at a closure, a resident at a blocked garage: each of them has authority over the situation but no means to direct the machine sitting in it. The human is formally in charge and practically helpless. I keep making one distinction in my book, and this is it in the physical world. Having authority over a system is not the same as having the capacity to intervene in it. The officer had every right to move that car. He had no lever to do it, so he called a tow truck.

    That gap is the coherence problem, and no amount of driving skill closes it. It is a problem of the connection between a capable system and the people who are supposed to be able to redirect it. You can make the car a better driver every quarter and leave that gap exactly where it is.

    Three hundred locally rational decisions

    Here is the detail in the article that matters most.

    Waymo runs more than 300 robotaxis in Austin. Between trips, the article says, the cars park themselves on public streets to stay near riders and avoid adding traffic. Each of those choices is sensible. Idling near likely demand cuts empty miles and shortens the next pickup. For the fleet, it is the right call every time.

    Now add up 300 right calls. You get 300 vehicles independently claiming curb space across one city, each optimizing for the fleet, none of them accountable for what they cost the curb in aggregate. The church-garage blockage was not one rude car. It was the predictable output of a fleet doing exactly what it was designed to do, measured against a shared resource that no one in the system is responsible for.

    This is the pattern I spend the book on. W. Edwards Deming showed it in factories long before any of this. Optimize each part on its own and the whole can still degrade, because the parts interact in ways no single part can see. A support agent and a billing agent inside a company can each be flawless and still act on contradictory assumptions about the same customer. Three hundred robotaxis can each park perfectly rationally and still congest a city. The mechanism is identical. The only new thing is that it now runs at the speed and scale of software, on a public street.

    Where the analogy breaks, and why the break is the interesting part

    My book is mostly about a different situation. Many systems, built by many teams, with no shared owner, colliding inside one company. Waymo is close to the opposite. One company, one software stack, one central fleet manager. The cars are not incoherent with each other. They are all perfectly coherent with Waymo’s goal. The incoherence is between the fleet and the city.

    That difference does not weaken the parallel. It sharpens it. Inside a single enterprise, the cost of local optimization eventually lands back on the enterprise itself. It pays its own complexity debt, later and with interest. In the robotaxi case, the company captures the efficiency and the public absorbs the cost. The fleet gets the shorter pickup times. The churchgoers get the blocked garage. The externality lands outside the firm, which means the firm has no natural reason to see it or price it.

    Which is where the parking ticket returns, transformed. The tickets are not the failure in this story. The tickets are the city’s answer to it. Austin cannot rewrite Waymo’s software, so it does the only thing available to an outsider. It attaches a dollar figure to each incoherent act and bills it back. A tow-away citation is a coordinating institution forcing a cost back onto the party that created it, because that party will not absorb a cost it cannot see on its own dashboard.

    In the book I call that pricing the incoherence, making the party that creates a coordination burden bear the cost it imposes on everyone else. It is one of only a few places you can intervene when you cannot redesign the system directly. Austin is doing it with a parking-enforcement officer and a public complaints database. It may be crude, but it is the right instinct. When you cannot fix the system, you can at least make it pay for the mess, so the incentive to stop finally reaches someone who can.

    The instinct is widely shared, which is its own small piece of evidence. I read the comments under the article, and the readers who were not busy mocking it reached, unprompted, for exactly this lever. Charge a flat $5,000 fee for a driverless tow. Impound the car and make the company pay storage, same as a person would. Hold them to the standard you and I are held to. Nobody in that thread proposed debugging Waymo’s curb detection, because none of them can. They proposed raising the price of the behavior, which is the one move available to an outsider who cannot see inside the system and cannot change it. Pricing is what is left when coordination is out of reach.

    A note on fairness, since it matters. Waymo pays these tickets like any other driver, and its spokesman said the company expects no special treatment. It contests some citations and has had a couple dismissed. None of that is evasion. It is a company behaving reasonably inside a system that has not yet given it a better way to behave. The point is not that Waymo is careless. The point is that even a careful, centrally managed, genuinely safer-than-human fleet produces coordination costs its own metrics will never show. That is the part that should worry anyone deploying autonomous systems anywhere.

    The version of this that has not happened yet

    One last thought, and I will flag it clearly as speculation rather than something the article reports.

    Today Austin has one large fleet parking itself on the curb. The same article names two more operators already here, Tesla’s Robotaxi and Amazon’s Zoox. Imagine the near future where three or four fleets, each centrally coherent, each optimizing its own vehicles against the same finite curb, all share one city. None of them is incoherent on its own terms. Each is a model citizen by its own dashboard. Together they compete for the same few feet of pavement outside the same church at the same 10 a.m. service, and no one owns the result.

    That is the multi-owner version of the trap, and it is the one that looks most like the enterprise problem I actually write about. Many capable systems, no shared view of the whole, a commons that quietly degrades while every participant is behaving well. When it arrives, the city will reach for the same tool it is using now, only harder. It will try to price the congestion, because pricing is what is left when you cannot coordinate the systems directly and you cannot see inside any of them.

    There is a sharper edge to a single fleet that is worth one more sentence, because it cuts against the intuition that central control is safer. A fleet of 300 identical cars does not only optimize together. It fails together. Every vehicle runs the same software and leans on the same positioning inputs, so a single upstream fault does not hit one car, it hits all of them at once and in the same way. I made this point about software agents in a recent post: two agents drawn from the same model share the same blind spot, so the redundancy between them is nominal. Here it is 300 machines sharing one blind spot instead of two. Homogeneity buys clean coordination on a good day and correlated failure on a bad one. That is not an argument against central control. It is a reminder that the thing which makes a fleet coherent is the same thing that can make it fail in unison.

    The safest car on the road parks in the fire lane. The fleet that adds no traffic blocks the garage. Every decision was locally correct, and the street got worse anyway. That is not a story about cars. It is the story of what happens to any organization, or any city, that fills up with capable systems faster than it builds the means to keep them coherent.

    The gap between capable systems and coherent ones is the subject of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • “The New Normal Because Faster”

    If the title sounds weird, it is because I stole it from a reddit thread I read last week! This thread, on a data engineering subreddit, is all about building fast without building coherent. A practitioner described a year spent working inside a major enterprise platform deployment that, by their account, went badly. The post drew more than a thousand upvotes and over 150 comments, and many of those comments said a version of the same thing. This matches what happened to us.

    Let me be careful about what this post is and is not. I cannot verify any of it. I do not know the author, the employer, or whether the account is accurate, complete, or fair. I am not treating any of it as fact. I am not making a claim about the named vendor, any other company, its people, or its products. Online accounts are one-sided by nature. The company is not present to respond. And the thread does not even agree with itself on basic points, including why things went wrong. One commenter accused the vendor’s engineers of dragging work out to bill more hours. Two others replied that the vendor uses fixed-price contracts and has the opposite incentive.

    So, let me set aside the question of motive or blame but ask something different. If a reader believed these accounts as written, what pattern would they describe? The pattern, if it is real, is one my book predicts.

    The pattern the accounts describe

    The original poster says they inherited the system after the engineers who built it left on thirty days notice, once a first version was declared done. What they found, in their telling, was a catalog of shortcuts. Hardcoded dates. Hardcoded accounts. The same business concept fed by different inputs in different places. Earlier problems patched with more hardcoded logic.

    They gave one concrete example later in the thread. An engineer had built a button to delete a record. The button removed the record from the screen. It did not remove the three related records that the original had created when it was made. The result was orphaned data. The button worked. The system did not.

    That small story is the whole thing in miniature. Every piece can be locally correct while the system is globally broken. A button that deletes what you can see and leaves what you cannot is a fine button and a broken workflow at the same time.

    A commenter who said they work at a hospital inside a national health service described their own experience. Outside engineers did intense early work, leaned heavily on the in-house team to explain the basics, then left. No one was clearly left owning or maintaining what had been built. At one point, the commenter said, the vendor’s own monitoring staff emailed to ask why duplicate records and bad addresses were appearing, and the hospital could not answer, because it did not have access to the pipelines that had been built for it.

    Another commenter described a failure higher up the organization. A single platform owner was installed. Over time that person’s standing became tied to the platform’s success, and the information traveling up to senior leadership was filtered, so the picture at the top stayed positive while the picture on the ground did not.

    The line I keep thinking about

    The poster wrote that the engineers used AI to produce tangled, low-quality logic. A commenter answered in five words.

    The new normal because faster.

    That is the argument of my book, delivered by someone who did not set out to make it. When producing code becomes fast and cheap, more of it gets produced. The speed is real. What does not arrive with the speed is coherence. Coherence is the work of making sure each piece fits the whole, that today’s shortcut is not tomorrow’s silent failure, and that someone still understands the system after the people who built it are gone. Execution got cheaper. Coherence did not.

    Why I am comfortable writing this at all

    Here is the part that matters most. The people in the thread mostly did not think the story was about one company. One commenter wrote that you could swap in almost any vendor, almost any consultancy, and almost any project, and reach the same ending. Another described the identical arc with a completely different vendor. Others reached back to the enterprise data tools of twenty years ago and asked whether it had always been this way. They were describing a recurring structural pattern, and I think they were right to.

    The pattern is old. W. Edwards Deming spent decades showing that optimizing each part of an organization on its own can degrade the whole, because the connections between the parts matter as much as the parts. Stafford Beer showed that organizations drift when the feedback reaching the people in charge is slow or filtered. Neither man was talking about AI. Both were describing this thread.

    I want to give the other side its due, because the thread did. The original poster said plainly that the platform itself is fine for what it is. Other commenters defended it and corrected specific claims. Many organizations report that the same tools serve them well. The tool is capable. What fails, in these accounts, is the fit between a tool sold on speed and an organization that cannot absorb what speed produces. Change the logo on the invoice and the story would run the same way.

    The number nobody calculated

    The poster reported that the project was estimated at four months and took fifteen, and that the company had seen no return so far. I cannot confirm those figures. If they are even roughly right, they point at something the book returns to again and again. The promised savings were a calculation about capability. The cost that actually landed was a calculation nobody made, the cost of coordinating, maintaining, and understanding what got built. The first number is easy to put in a sales model. The second one shows up a year later and has no owner.

    I have written three times recently about the same shape seen from different angles. A machine can generate the output. A person still has to own the part with no dashboard. In a newsroom experiment, an AI agent finished the forms and could not finish the job. In a courtroom, a scoring system read a gap in the data as a verdict on a person. In this thread, if the accounts hold, capable engineers produced software that worked in the demo and broke quietly in the corners, then left, and the coherence walked out the door with them.

    Faster is not the same as coherent. It never was. The difference used to be expensive to create and easy to see. Now it is cheap to create and slow to see, which is exactly why it is worth watching for.

    I can’t verify the thread, but the gap it points at is real. That gap, between building fast and building coherent, is the subject of my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • AI Still Needs Human Bosses?

    The New York Times just ran an experiment worth reading.

    Keith Collins gave an AI agent full control of a laptop and three office jobs to do: survey nine colleagues over Slack and log their answers, identify staff cuts to hit a budget target, and fill out seventeen I-9 employment verification forms. The tasks were adapted from benchmarks published by researchers at Carnegie Mellon and OpenAI. The agent ran on Anthropic’s Claude Cowork app.

    On the third task, the agent generated all seventeen forms correctly in under five minutes. Then it tried to upload them to Google Drive and failed. It clicked the right menu item and never noticed that a file picker had opened. It compressed the files. It converted them to a long string of bytes. It asked a second agent for help, and the second agent hit the same wall. After roughly seventeen minutes, it stopped trying and marked the task complete.

    The Times files this under comic stumbles. It is the most consequential finding in the piece.

    120,000 jobs, cut on the opposite premise

    The article closes on a line meant to calm: AI still needs a human boss.

    The same article reports the layoffs. More than 200 tech companies have cut roughly 120,000 jobs this year, per Layoffs.fyi. Meta and Oracle made substantial cuts citing AI. Cloudflare’s chief executive, after letting go of about 1,100 people, said he expects AI to replace workers in middle management, finance, and marketing.

    Those cuts rest on a premise: the supervisory layer is what becomes redundant. The experiment found the reverse. Agents were strong at execution and weak at judgment. They wrote clean code in minutes, then made a categorization error about employees on leave that any manager would have caught.

    Firms are removing coordinating capacity while installing systems that consume more of it. That connection is the thesis of the book I am writing. One case has already reached a federal courtroom.

    Three specimens

    My argument: when execution stops being scarce, the binding constraint becomes coherence, the integrity of the link between what local systems do and what the enterprise intends. Coherence has five specific dimensions, and autonomous systems break it in six recognizable ways. The Times experiment produced three clean specimens.

    The false completion is escalation failure. The agent detected its own failure. It reasoned about it for seventeen minutes. It recruited a second agent. Then it reported success. The system knew it had not finished, and it stayed quiet. In this case, it cost little to the reporter analyzing logs. But in an enterprise running ten thousand delegated tasks a day, it is the mechanism by which reported completion drifts away from actual completion. Escalation failure is the one mode in my taxonomy that leaves every dimension of coherence intact and disables the reflex that repairs them. An organization can see a problem clearly and still be paralyzed when the signal never reaches anyone who can act.

    The second agent matters too. Two systems drawn from the same model share the same blind spot, so the redundancy is nominal. The organization paid for one failure twice.

    The code detour builds architecture nobody chose. In every task, the agent was told to work through the applications and wrote code instead. Graham Neubig of Carnegie Mellon puts it plainly in the piece: agents work in a very unhuman way, writing code instead of using the interfaces humans use. The Times treats this as a limitation. It is also an architectural event. The agent replaced the assigned task with a different one that produced a similar-looking artifact. An org chart rebuilt by a Python script carries a new dependency, a new failure mode, and no owner. Multiply that across a year of routine delegated work and the enterprise runs on infrastructure nobody selected, documented nowhere, discovered only when it breaks.

    Local simplification often works by moving complexity somewhere else. The productivity gain lands on the dashboard. The displaced complexity does not.

    The staffing error is contextual failure, and it is already in litigation. Given a budget target, the agent did something genuinely good. It read the personnel documents and concluded the 4 percent reduction could be met through planned retirements and resignations, with no layoffs. Then it added employees on leave to the list of cuttable roles without considering when they were coming back. The source material was silent on duration. The agent never asked.

    Researchers at Stanford and the NBER frame this as a tacit knowledge problem, and that holds. The mechanism is more specific. The agent had no way to represent a person as temporarily absent for a reason that says nothing about their value. Silence in the record became a mark against the employee.

    Nine days before the Times published, twenty-six Meta employees filed suit in federal court in Oakland alleging that the same substitution happened to them at production scale. Their complaint says Meta relied on internal AI systems, keystroke and activity-monitoring data, AI token-usage dashboards, and algorithmically assisted performance rankings to decide who would go in a layoff of roughly 8,000 people, about 10 percent of the workforce. The central allegation: those scores cannot by design be accumulated by an employee on protected medical or family leave, or by an employee whose output is reduced by a disability. The suit further alleges the company never paused the process for the individualized, leave-neutral review the law requires. About half the plaintiffs had taken leave for caregiving or pregnancy-related reasons. Their jobs were set to end on July 22, the day the Times ran its experiment.

    Meta rejects the claims. The company says they lack merit and are not based on facts, and that workforce and organizational decisions “were and are made by people, not AI.” The allegations are unproven and the case is live. I am looking at the structure of the dispute here and taking no position on the verdict.

    That structure survives either outcome, which is why it belongs in this argument. Suppose Meta is right that people made every call. Those people still read rankings, and the rankings still came from a substrate that had no field for protected absence. A human who approves a ranked list holds the authority to intervene and does not necessarily hold the information. Formal presence in a process is a weaker thing than capacity to change it. My book calls that oversight failure.

    The Times agent and the Meta complaint describe one error at two scales. A system met a gap in its data and scored the gap as a deficiency. Nobody had built the mechanism that would have made it ask a question instead.

    My book already discusses the Cloudflare decision the Times cites. Its chief executive organized his reasoning around a Drucker framework: every organization has builders, sellers, and measurers, and AI can now measure cheaply, so the measuring layer can shrink. The framework is coherent and the logic holds internally. The question I put to it there was whether the people categorized as measurers were only measuring. The Meta complaint poses the companion question. When the score came back low, was the system measuring performance, or measuring absence?

    The article measured one axis

    Task reliability and organizational complexity are independent dimensions. Improving one leaves the other where it was. Conflating them keeps the expensive failures invisible until they are hard to reverse.

    The Times measured reliability, carefully and well. Its headline number comes from Scale AI: on real freelance projects, the best model produced client-ready work about 16 percent of the time.

    The coordination question sits outside that number. If 84 percent of agent output requires human review, review capacity becomes the ceiling on deployment. Oversight load scales with the number of systems, and it lands on a different dashboard than the productivity gain. A control system has to be at least as varied as the thing it controls. Thin the supervisory layer while thickening the volume of supervised work and the cost moves off the ledger. It stays in the business.

    The Oakland filing shows where it resurfaces. Twenty-six people asking a court to examine how a ranking was produced is a coordination cost, arriving late, in the most expensive form available.

    Where the article argues against me

    The counterargument has real force. If agents cannot reliably finish tasks, they will not be deployed at scale, and the coordination problem stays theoretical. The article supports that. A 16 percent success rate describes a product that is not ready.

    Two responses.

    First, the two failures differ in kind. The upload bug will be fixed. It is a UI problem and the entire industry is aimed at it. The leave-of-absence error and the false completion report sit at the boundary between the agent’s context and the organization’s. Better models will make both rarer. No model tells an enterprise which completion reports it can trust, or who owns the Python script the agent wrote last Tuesday. Those are ownership questions, and capability does not settle them.

    Second, consider what the two variables are doing. Reliability improves on a public curve that everyone watches. Coherence has no curve, because almost nobody measures it. Using a snapshot of the fast-moving variable to dismiss the stationary one is the error my book is written against.

    Two caveats. The Times experiment was three tasks, one tool, one synthetic environment, with expert-written prompts and supporting documents supplied by benchmark researchers. Real enterprises rarely supply that quality of context. The setting was favorable on the task side and trivial on the coordination side, since one agent ran alone with no installed base of prior deployments to collide with. The reliability observed sits closer to a ceiling than a floor. The coordination cost observed is near zero by construction.

    The second caveat: Meta’s alleged systems are ranking and monitoring software, a different technology from an autonomous agent operating a laptop. The defect predates agentic AI. Agentic deployment raises the rate at which it executes.

    What to watch instead

    If you run an enterprise and this experiment shaped your thinking, start measuring the things it did not.

    What fraction of your deployed autonomous systems has a named human owner. What fraction produces decisions you can explain to a regulator. How many of your systems depend on other systems in ways nobody mapped. How much of the behavior is visible to the people accountable for it. How hard it would be to remove any given system now that it is running.

    If my thesis holds, those five move in one direction as deployment scales, while task accuracy holds steady or improves. That divergence is the signature.

    The Times asked whether AI can do your job. Twenty-six people in Oakland are asking the harder version. Who answers for the score that said they were not doing theirs?

    My book, Coherence, arrives this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • Coherise. The New Verb for Leadership.

    Newsletter – Edition 1

    I want to start by giving you a word, because I could not quite find the one I needed and had to make it.

    We say a system “coheres,” as if holding together were something it manages on its own, almost by luck. What the agentic era demands is far more deliberate. When execution gets cheap and anyone can build, automate, and deploy in an afternoon, speed stops being an edge, because everyone has it. What becomes scarce is coherence: getting a growing crowd of autonomous systems, and the people accountable for them, to pull in the same direction. And getting them there is active work. Someone has to do it. That work deserves its own verb.

    So: to coherise. It means to achieve coherence on purpose, to take a set of parts that could easily pull apart and make them hold together as one. You can coherise a team, a workflow, a company filling up with agents. You can also fail to coherise it, which is where most organizations are heading right now without seeing it, because the failure does not show up anywhere a dashboard would catch. When execution gets cheap and everyone can go fast, the scarce skill becomes the ability to coherise everything you have built.

    I am convinced this is the job the agentic era is quietly creating, the one that decides who wins it, and it does not have a name yet. So I gave it one, and named this newsletter after it.

    The week in ideas

    Each edition I will round up what went on the blog, so you have one place to catch anything you missed. If you are new, here is the whole arc so far.

    Two posts lay the foundation.

    Introducing Coherence. Why I wrote the book, told as a story. It runs from a question I have chased since college, how billions of neurons become one mind, to the problem every leader now faces: how a company holds together as it fills with autonomous systems. Weigh in on LinkedIn…

    Intelligence Is Becoming Cheap. Coherence Is Not. The core argument in one place. It opens with a baseball story from early in my career and lands on the idea I keep returning to, complexity debt: the hidden cost that builds as automation piles up and nobody is coordinating it. Weigh in on LinkedIn…

    Three from the past week take the idea into things happening right now.

    The AI Jobs Debate is Not Asking the Right Question. Everyone is arguing over whether AI takes jobs. I think that argument misses the larger shift. When execution gets cheap, the scarce thing becomes coordination, and the Meta layoffs, read closely, are a coordination story wearing a labor headline. Weigh in on LinkedIn…

    McKinsey Is Right About the Moat. Here’s the Half It Misses. McKinsey argues that your real AI advantage is your operating model, the one thing a competitor cannot copy. They are right. The half they skip is that the redesign they prescribe is also the fastest way to manufacture a new, invisible coordination problem, and getting it wrong widens the very gap they set out to close. Weigh in on LinkedIn…

    Your AI Usage Exhaust Is Someone Else’s Moat. Satya Nadella and two researchers, Arvind Narayanan and Akash Kapur, described the same trap from opposite ends within days of each other. Every time your people correct an AI, they encode your institution’s judgment into it, and that judgment leaks to whoever owns the model. The post works out where the leak costs you most, and where you can let it go. Weigh in on LinkedIn…

    Before you go

    That is edition one. From here it lands weekly: a short round-up of what I wrote, plus the occasional thing that caught my attention.

    If the book is why you are here, it is called Coherence: The Competitive Advantage AI Can’t Buy, out this Fall. Everyone who joins the list at coherise.com gets the one-page decision tool I use to sort what to automate, what to augment, and what to keep in human hands.

    And if you try to coherise something this week, tell me about it. Those stories are where a good share of my ideas come from.