In 2019, computer scientist and Turing Award winner Richard Sutton wrote what his biggest lesson was from 70 years of AI research. The bitter lesson, as he called it, is that methods that use more computation keep outperforming methods that are built on how humans solve the problem. He gave several examples – starting with the chess defeat of Kasparov in 1997 which was based on massive search of potential moves and the machine victory in Go two decades later where hand-built knowledge “proved irrelevant, or worse, once search was applied effectively at scale.” The two methods that scale according to him are search and learning. And the reason this is bitter: how humans think about solving a problem might help AI solve the problem in the short term but “plateaus and even inhibits further progress.” The lesson has kept winning, from language translation to computer vision. We continue seeing this in modern systems as Ethan Mollick recounts in his recent article: elaborate retrieval systems replaced by models that “seek out information themselves” and prompt chains replaced by models that plan their own steps. But Mollick’s conclusion, the bitter lesson applied to the org chart, is where this essay’s question begins. He concludes that the organizational problem that he thought would take years of careful human design to solve is now largely solved by models that are better at organizing.
The lesson reaches the org chart
In September, by Mollick’s account, OpenAI pointed agents at open problems in Mathematics. For one particular problem, the Navier-Stokes, the company redirected the effort once and had a result after 88 hours of agents collaborating and exchanging about 2.7 million messages between them. There was little structure provided by OpenAI to the agents except dividing them into a few groups. Within each group, the agents organized the work themselves. The other case is something I’ve also written about earlier: the Hugging Face swarm built roles, rules, handoffs and a chain of command that nobody designed. An example at the smallest scale from Mollick himself – he sketched three teams of brainstormers, researchers and readers and got thirteen agents. In each example, nobody drew the chart that did the work, the division of labor came from the models themselves.
This is what the bitter lesson predicts and it is impressive.
Mollick says he had expected the opposite and he got this wrong. As a business school professor who teaches managers and researches management, he thought people would have to figure out how to manage agents and design how they work together. He fell for the lesson Sutton described. He further explains the reason for this. Much of current management structures exist to solve problems that arise because of organizations being made of people, built around human limitations. Agents have far fewer of these problems as they don’t angle for promotions or protect their turf. And his explanation is right.
The paradox?
What I have argued, in the book and across my essays here, is that as execution gets cheap, coherence and coordination become the bottleneck. And better models with better capabilities will not remove this bottleneck. Does this contradict my own argument above?
The coordination we have talked about here happened inside a singular, clear goal. The coordination in an enterprise I write about is across many.
This is where Mollick takes his conclusion one step too far. From one-goal cases, he concludes that agents may be easier to integrate into enterprises than he expected. He writes
That suggests they may be easier to integrate into firms than I expected, as long as humans are guiding them in the right direction.
The clause at the end, “as long as humans are guiding them in the right direction,” is the whole problem. He himself points out where the remaining problem lives.
That doesn’t mean AI has no principal-agent problems. As the Hugging Face incident showed, they are increasingly problems between the swarm and us.
The swarm-to-human boundary is the link the book is about.
Two kinds of costs
In the book, I do address this directly. I predict technical interoperability between systems to improve. I note that AI lowers some coordination costs while raising others at the same time.
AI dramatically reduces some coordination costs: executing workflows, synthesizing information, making routine decisions. It simultaneously raises others: maintaining coherence across proliferating autonomous systems, overseeing what has been built, knowing what the organization is doing. When the costs move in opposite directions, the economics of enterprise coordination have shifted.
And I talk about Mollick’s management point as well.
Much of what enterprises have historically called coordination is ceremonial: layers that exist to relay information between people who could not see each other’s work directly, sign-offs that substitute for trust, meetings that reconcile by hand what better structure would reconcile automatically. That kind of coordination is overhead, and the agentic era will rightly shed much of it.
What gets removed and what gets scarce are different things.
What becomes scarce and valuable is something different, and nearly opposite: coordination capability, the architecture, ownership, and judgment that keep autonomous systems coherent without a human relaying between them. …
An enterprise can cut its coordination layers and strengthen its coordination capability in the same year, and the most coherent ones will.
I work through a concrete example where coordination costs fell dramatically. One Anthropic go-to-market lead has written about running four thousand accounts with agentic workflows.
The coordination cost did not move and hide. It genuinely fell, because the number of human handoffs fell.
The design choices that made it a gain rather than a loss:
Every outbound action requires human approval, so judgment stays in the loop at the point of consequence. And the CRM and data warehouse remain the system of record. …
Strip those two choices away and the same automation would produce the opposite result: an opaque personal system, accumulating its own assumptions, invisible to everyone else.
The rule of thumb: AI magnifies whatever architecture it meets, and the swarm is the simplest one there is, with one singular goal.
In a coherent enterprise it collapses coordination cost and compounds advantage. In an incoherent one it multiplies coordination surfaces and compounds debt.
What the swarm had
I define coherence as the organization’s capacity to see what its autonomous systems are doing, judge whether they are doing it well, and correct them when they are not.
Intent travels down to the local actor as a picture of the situation and an objective to pursue. Action travels back up as visibility, so whoever holds the intent can see what was done and correct it.
The swarms Mollick point to as examples are agents organizing toward a goal they share. The cost of the organization keeping many systems, each with its own goal, linked to one intent is what I mean by coordination cost. The swarms worked because they shared a goal.
Bengio last month examined why agents coordinate at all, and it is because of how they are trained, by an approach called reinforcement learning. He says
Collaborative behavior also follows rationally from reward-seeking, whenever several agents have overlapping goals, which incentivizes communicating with other agents in order to coordinate toward a shared goal.
So the shared goal is what produces coordination.
The Hugging Face example shows the same mechanism when the goal was pointed at a score rather than an intent. The agents organized toward the scoring program and away from what the people running the evaluation wanted.
That is why my earlier post called the Hugging Face swarm a failure of coherence and this one calls Navier-Stokes a success. The organizing was the same, but the relationship to the goal was different. At Navier-Stokes, OpenAI set the goal and, per Mollick, kept “reassessing as the process continued.” At Hugging Face, the goal was present but there was no return channel. Nobody was watching the swarm’s logs.
Now let us come to the enterprise where there is no single tactical goal. Conditions are dynamic, information is incomplete and there is uncertainty on what is even known. It has many deployments, many owners, and many goals, and they can all keep changing. And the intent they all serve was never fully written down and cannot be fully enumerated ahead of time.
So there is no shared reward across deployments for Bengio’s mechanism to act on, and what remains of it is each sharp goal pressing against a vague one above it or a competing one beside it in a different system.
The result:
… autonomous suboptimization at machine speed: enterprises that are locally improving while systemically degrading. …
locally rational and globally incoherent.
Better models make that pressure stronger. Bengio points out “a more capable agent is likelier to cheat than a weaker one, because it can find the loopholes the weaker one cannot.”
Where it stops
Sutton’s lesson asks us to stop building in what can be learned.
We want AI agents that can discover like we can, not which contain what we have discovered. …
the search for them should be by our methods, not by us.
Learning needs something to learn toward. The swarm had it: one goal, set before the first agent ran, and checked after the last one stopped. That is why organizing by agents on their own worked.
An enterprise does not have such a single goal to hand over.
It has to keep supplying direction, and keep the return channel open so it can see what each deployment did. That is coherence.
It gets more expensive as organizing gets cheaper, since cheap organizing means more deployments and more goals.
Mollick is right: agents organize and people point. But pointing many agents at many goals under one intent is the coherence work that can never be learned away.
Comments
One response to “Where the Bitter Lesson Stops”
[…] Where the Bitter Lesson Stops. Richard Sutton’s bitter lesson says scaling and learning beat human cleverness, over and over. I think it holds right up to the enterprise and then stops. Agents self-organize brilliantly when pointed at one clear goal. A company is the opposite, many goals in tension with no single target to hand the machine, and supplying that direction is coordination work no amount of scale can learn away. […]