Essays on how companies hold together as they fill with abundant intelligence. Written for CEOs, enterprise leaders and executives navigating the organizational impact of AI and agentic AI.

AI and agentic systems are changing organizational design, operating models and the way companies work. The problem isn’t simply redesigning the organization for AI. It’s keeping the redesigned organization coherent as AI accelerates complexity.

Looking for something else? My academic publications and my Concentric AI writing live elsewhere.

Can Frontier AI Outdo MBAs?

Three of the top business schools in the country just tested frontier AI on the analytical work their MBAs are trained to do, and the models scored in the high eighties. If you run a company, that number is coming for you soon, probably in a deck that recommends cutting a layer of analysts. So it is worth being exact about what it measures, because the exact answer is more useful, and more limited, than the headline.

The paper is BusinessCaseBench, from researchers at Wharton, Carnegie Mellon, and Harvard Business School. They drew 615 questions from real business school cases across eighteen disciplines, from strategy and finance to leadership and ethics. Each model read a case cold and wrote its analysis. A separate grader then compared that analysis against the reference solution the instructor had written, item by item. The models never saw the reference. As far as the model was concerned, it faced an open business question with no answer attached, and it answered well. Under the main metric, Claude Sonnet 4.6 covered about 88 percent of what the instructor’s solution contained, and GPT-5.4 about 87.

The model was not helped by the answer key. It could not see the rubric, did not use it, and produced its analysis from the case alone, the way a consultant works from a brief. These were, from the model’s side, genuinely open questions. So it is fair to expect that a model which writes strong analyses on 615 unseen cases will write a strong analysis on the 616th, which is your real one.

The problem is that “the model will write a similar analysis” and “you will get a similar result” are different claims, and only the first one is what the benchmark tested.

A score is a comparison

“The model scored 88 percent” is not a fact about the model’s answer by itself. It is a fact about that answer measured against a standard. Two things had to exist for the number to exist: the analysis the model wrote, and the instructor’s solution it was checked against. The score is the relationship between them.

Now move that setup to your company. The model can still write the analysis. But the standard it gets checked against does not exist. The 88 was a statement about how well the answer matched a known-good answer.

This is not word games. It is the difference between “the model is competent” and “you can trust the output.” The first is about the answer. The second is about checking the answer, and checking requires a standard. The benchmark supplied the standard. Your hardest decisions do not.

Knowable-but-hidden is not the same as unknowable

The tempting reply is that the real world is just the benchmark with the answer hidden. The model handled hidden answers fine; a real decision is one more hidden answer.

But the benchmark’s answers were not hidden. They were knowable in the first place. The professor had already worked the case. A correct answer existed; the model simply was not shown it. That is a different situation from the one you are in when you decide whether to enter a market or restructure a division. There, no correct answer exists yet. It has not been written by anyone, because the outcome that would settle it is years away, the criteria for “good” are contested by the people in the room, and there is no counterfactual to check the decision against even after the fact.

A graded case has a knowable answer the model didn’t see. A live strategic decision has an unknowable answer that does not exist to be seen. A student who scores 88 on a past exam she took blind will likely score about 88 on the next past exam. But it does not follow that she will make good venture bets, even though both feel like hard open-ended judgment, because a venture bet has no marking scheme, then or later. The model is the student. The benchmark is the past exam. Your boardroom is the venture bet.

So the benchmark is strong evidence for a real claim: frontier models are good at producing structured business analysis, and getting better fast. It is not evidence for the claim that the score predicts a good outcome on decisions whose standard has to be invented rather than looked up. Inventing that standard, deciding what a good answer to your actual question would even need to contain, is the judgment.

That argument stands even if the models are excellent.

The model is fully right about half the time

The researchers scored the answers two ways. The headline 88 percent is partial credit: how much of the instructor’s checklist each answer covered. Then they ran a stricter count. On how many questions did the answer satisfy the entire checklist, every item, no gaps? That fell to roughly half.

On cases where a correct answer was knowable, the leading model produced a complete answer about one time in two. The authors put the point in their own title for that result: these are drafts, not verdicts.

Now extrapolate honestly. If you carry this model into the wild and expect “similar performance,” you are also carrying the incompleteness. The typical output is strong and missing something at the same time. In the study, a human holding the instructor’s solution caught the missing half. In your firm, if you removed the person who could catch it, the missing half is still missing and nothing catches it.

Grader didn’t check for what shouldn’t be there

The grader verifies whether each expected point is present. By construction, it does not scan the answer for confident, invented, or wrong material that sits outside the checklist. A response can hit the expected points and also assert three plausible fabrications and still score well, because nothing in the method is looking for the fabrications.

In the study that blind spot is harmless, because a grader with the solution ignores the extra material. In your company it is the whole risk. The fabricated line rides along inside a well-organized analysis, unflagged, and the reader who could catch it is the domain expert the strong score seemed to make optional.

Why “drafts, not verdicts” is the expensive finding

A draft that is 88 percent right sounds like a bargain, and sometimes it is. But it moves the work rather than removing it. Obvious garbage you would catch. Verified truth you could trust. The strong, incomplete, possibly-embellished draft is the expensive case, because it earns your trust on the parts you can see and hides its gaps in the part you would have to already know the answer to find. Catching what it left out requires someone who knows what a complete answer contains, which is exactly the expertise the draft appeared to retire.

So the generation got cheap and the checking did not, and on this kind of work the checking cannot be sampled down. You cannot review only the flawed answers, because the flawed ones are the ones that look fine. You review all of them, or you review none and call it oversight. An organization that reads the 88 as license to remove the reviewers has not automated the analysis. It has removed its own ability to tell when the analysis is wrong.

What it means if you run something

Read correctly, the benchmark is good news. Frontier models are genuinely strong at producing structured business analysis, and that is real leverage you should use. The error is reading a score-against-a-known-standard as a readiness-to-deploy-without-a-standard.

Two questions to ask of any AI you are about to trust with judgment work.

Does this task come with a knowable answer, or is defining the answer the actual job? Where a defensible right answer exists, straightforward analysis with a standard you could write down, let the model run, and expect it to perform in the wild about as it did on the bench. That extrapolation is fair, and worth taking. Where the standard itself has to be invented and argued, the model is drafting, and a person still owns the decision.

And who checks the drafts, and do they know enough to catch what a confident draft leaves out? If the answer is nobody, or nobody who could tell, the high score is measuring something you will not actually get.

The models are good at the framed question. Your hardest problems arrive unframed. The competitive advantage was never in answering the case. It was in knowing which case you were actually in.

This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.