Essays on how companies hold together as they fill with abundant intelligence. Written for CEOs, enterprise leaders and executives navigating the organizational impact of AI and agentic AI.

AI and agentic systems are changing organizational design, operating models and the way companies work. The problem isn’t simply redesigning the organization for AI. It’s keeping the redesigned organization coherent as AI accelerates complexity.

Looking for something else? My academic publications and my Concentric AI writing live elsewhere.

220% More Code, 40% More Incidents

The Wall Street Journal reported in late July that corporate America was pulling back on AI token spending. “Tokenmaxxing” had a scoreboard that measured spend. Nobody kept track of what it cost to absorb. Meta ran the experiment at large scale, alongside a restructuring plan it called “Project OT,” and reporting in Reuters has made its numbers public. Quoting an internal post by the CTO in early June, Reuters reported an increase of 220% in code changes to internal platforms. Major technical and security incidents spiked 40% from the previous year, with the time employees had to spend firefighting them up 70%. Katie Paul, the reporter who broke the story, in a public Q&A on Sep 3, said that “tokenmaxxing pressure is out” now and the guidance is to use AI where it makes sense.

The verdict from Zuckerberg himself was that they went too far ahead of where the technology was and that the “trajectory of the agentic development over at least the last four months hasn’t really accelerated in the way that we expected.” I spent almost a decade and a half building AI inside large companies before cofounding a data security company, and this sequence is familiar. While the verdict is partially correct, it teaches the wrong lesson. Zuckerberg told employees in April, explaining the rationale for the layoffs, that there were two major cost centers: “compute and infrastructure” and “people related things” and that investing more in one would leave less capital to be allocated to the other. It is true as a budget line. But it is wrong as a description of how AI output gets absorbed. Each output from AI has to be reviewed, integrated, and when it breaks, repaired. That is work that has to be done by people until the organization builds structure to do it: named owners for each system, constraints to eliminate entire classes of errors, walls between systems so they don’t contaminate each other, and escalation paths that are actually answered. The capacity that work requires depends on how much is deployed and how the systems interact, not on headcount. The 220% increase in internal platform code changes is one manifestation of the deployment surface. 

The causal claim comes from Meta itself. Reuters reports that as early as March, infrastructure teams flagged reliability warning signs from the surge in AI-written code. An internal post in April said that unchecked agents were taking “large-scale, disruptive actions that humans are unlikely to execute.” Changes that reached users as new or improved features rose 36%, against 220% increase in code. If that implies much of what was produced was noise, it still had to be read and reviewed to be recognized. Review load rises with the volume and not with value. As the Reuters reporter recounts in the Q&A, AI-caused incidents were “weirder and less predictable” than human-caused ones and which is how 40% more incidents became 70% more firefighting. The signal was present in pieces: on the reliability dashboard, in the on-call rotation and in the March and April posts. But no one collated it against the AI spend. 

EY put the principle in language that boards can understand: every organization has a ratio of builders to overseers, and governance capacity has to match building activity. Meta’s Project OT inverted it. While the code output increased, teams traditionally composed of 10 to 20 were restructured into pods of 3 to 5, and about 8000 people were laid off in May. What is not clear if there was any other structure in place to absorb what the smaller teams could not. Each new system adds interactions to every existing system. Coordination load rises with combinations and not with headcount. Meta is not the first to pay the price. Ford hired 350 veteran engineers to catch quality problems its automated systems missed. And IBM is tripling entry-level hiring because, in its HR chief’s words, without a pipeline “the well simply dries up.” Reducing token consumption or “thrift-maxxing” does not settle this bill either. Cheaper models mean more models, more routing, more repair, and more coordination. 

Meta says Project OT was a scenario exercise and not every scenario would run. Some of the rise in incidents reflects more usage and Meta has declined to comment on the figures. Zuckerberg’s timing admission is more candor than most CEOs admit, and tracking incident numbers at all is more than most companies do. And he still expects models to catch up within the next three to six months. But that is the point – better models raise the volume and the absorption bill grows with it. 

Before approving an AI-justified headcount plan, boards should ask for three numbers. What happens to review and incident volume when output rises 50%? Who owns each deployed system when the people who built them move on? What is the ratio of builders to overseers, and how is it chosen? Then ask for the incident line next to the token line, and ask what is meant to absorb the output once the people doing it by hand are gone. If the answer is nothing, the plan is a bet that absorption is free. Size the absorption capacity by how much is deployed and not how many are employed. Meta’s own correction points in the same direction: no more company-wide layoffs this year, selective retention packages for people it was about to lose, and the usage mandate withdrawn. That is Project OT run in reverse. 

Zuckerberg told his staff the company has two cost centers, compute infrastructure, and people. Meta’s own experience shows a third: the cost of getting what the first produces through the second. While it appeared nowhere in the budget, it showed up everywhere in incident logs. The next plan that treats those two as competing should be asked to show the third one first. 

Comments

One response to “220% More Code, 40% More Incidents”

  1. […] 220% More Code, 40% More Incidents. Meta ran the experiment at scale. Under a plan called Project OT, it restructured teams of 10 to 20 into pods of 3 to 5 and laid off about 8,000 people. Its own internal numbers, surfaced by Reuters, tell the rest. Code changes to internal platforms rose 220%. Major technical and security incidents rose 40%. Time spent firefighting them rose 70%. Only about a third of the new code reached users as features, so much of it was noise that still had to be read and reviewed. Zuckerberg’s takeaway, that they moved ahead of the technology, is partly right and teaches the wrong lesson. AI output has to be reviewed, integrated, and repaired, and that absorption capacity scales with how much you deploy and how the systems interact, not with headcount. There is a third cost center beyond compute and people: the cost of getting what the machines produce through the people who remain. It never appeared in the budget. It showed up everywhere in the incident logs. […]