A digression. This blog is usually about enterprises. This one is about mathematics, and about a question I have been fascinated with.
In May, an OpenAI model disproved a conjecture that Paul Erdős posed in 1946.
The problem is easy to picture, which is part of its charm. Scatter some dots on a page. Count the pairs that sit exactly one inch apart. As you add more dots, how fast can that count grow? Erdős guessed there was a ceiling, and that a grid-like arrangement came close to it. For decades most mathematicians thought he had it right. The model found arrangements that beat the ceiling, and kept beating it as the number of dots grew without limit.
The surprising part was the route. The unit distance problem belongs to discrete geometry. The solution came through algebraic number theory, which studies something else entirely. Writing in the Wall Street Journal last month, the statistician Daniel Kipnis makes an observation about this that I keep returning to. Cross-disciplinary borrowing is not new in mathematics. Descartes did it in the seventeenth century. What has changed is scale. A mathematician can spend an entire career in discrete geometry and never acquire the tools of algebraic number theory, because a human career is short and those tools take years. A machine has no such constraint. Its reach across the field is bounded by the cost of computation and nothing else.
So the machine went somewhere no specialist would have thought to look, and came back with a counterexample.
Then Kipnis asks the question that makes this interesting. If nobody understands a proof, is it a proof at all?
Mathematics as a social act
He gets his answer from Reuben Hersh, who spent a career arguing that mathematics is a social phenomenon rather than a collection of eternal truths sitting somewhere waiting to be found. On that view, a mathematical fact does not become part of mathematics by being true. It becomes part of mathematics by being discovered, explained, and absorbed into what the community understands. Progress is a form of communication. A proof explained badly does no more good than a proof that is wrong.
Kipnis notes that OpenAI seems to have understood this instinctively. It did not publish the output and walk away. It worked with prominent mathematicians who verified the argument and wrote a companion paper making it intelligible to the field. Without that second step, the result would have occupied a strange position. It might have been true. But if no human could confirm it or follow it, what would its truth consist of? On Hersh’s account it would not yet be a mathematical result.
This is a strong claim and I find it persuasive. It also has a problem, which two other people, in another discussion, illuminate.
Is understanding a crutch?
On Quanta’s podcast The Joy of Why, Steven Strogatz recently interviewed Lauren Williams, the Harvard mathematician who helped start the First Proof project. Afterward Strogatz and his co-host Janna Levin, an astrophysicist, kept talking, and the conversation turned to something more unsettling than job displacement.
Strogatz asked whether beauty will still guide mathematics once machines are doing it alongside us. Beauty in the working sense, meaning the aesthetic pull that tells a mathematician which question is worth asking and whether an argument is on the right track. Earlier in the discussion, Strogatz and Williams had discussed about her philosophy on beauty: “If you ask a question and the answer is not beautiful, that means you asked the wrong question.”
Levin’s answer is the best thing I have read on this in months. One of the things beauty does, she said, is make a complicated subject comprehensible. Then she gave the reason she needs that: “I don’t have infinite compute.”
Elegance is not decoration. It is compression. Understanding is the technique a bounded mind uses to fit something enormous into a space the size of a human head. We prize proofs that are short, surprising, and clean because we cannot hold the long ugly ones. As Schmidhuber, a leading AI scientist, explains, a computationally limited observer finds something simpler and more beautiful once she learns to predict and compress the data in a better way. Herbert Simon spent a career making a version of this argument about organizations, which exist in part because no individual can hold the whole problem, so the problem gets cut into pieces a person can carry. Mathematical understanding looks like the same adaptation, running on the same constraint.
Which invites the obvious follow-up, and Strogatz asked it. If understanding is a workaround for our limits, is it overrated? He suggested we might be confusing means with ends. If the goal is true theorems and reliable prediction, comprehension is the ladder, and once you are up you can kick it away. He offered a medical analogy. If a therapy saves a life, you may take it without understanding why it works.
He also gave the other side its due, which is that some people regard science without understanding as a diminished thing, and he said he could see both positions.
Where the analogy breaks
But notice what the medical case is quietly relying on.
You can accept a treatment you do not understand because you have another way of knowing it works. The trial. The outcome is observable, the effect is measurable, and the verification runs on a completely separate track from the explanation. Understanding is genuinely optional there, because something else is doing the job that understanding would otherwise do.
Mathematics has no second track. There is no experiment that shows a theorem is true. You cannot run a trial on a conjecture. The only instrument the field has ever had for establishing that a statement holds is a proof, and a proof is a piece of writing addressed to another mind. In mathematics, verification and explanation are not two activities that happen to co-occur. They are the same act.
That is why Hersh’s position is stronger than it first appears, and why “understanding may be overrated” does not transfer cleanly from medicine to mathematics. Give up on understanding a proof and you have not traded comprehension for reliability. You have given up your only method of knowing.
The escape hatch, and what it costs
There is one way out, and it is real. Machines can check proofs.
This is not new and the mathematics community has been living with the discomfort for fifty years. The four color theorem fell in 1976 to an argument that included computer case-checking no human could reproduce by hand, and mathematicians argued about whether that counted.
The sharper case is Thomas Hales. In 1998 he announced a proof of the Kepler conjecture, about the densest way to stack spheres. The Annals of Mathematics assigned twelve referees. They worked for four years. They concluded they were ninety-nine percent certain the proof was correct, and admitted they could not independently verify the thousands of lines of computer code it rested on. Full publication came nearly eight years after submission. In a retrospective written years later, Hales says plainly that the review dragged on until the referees became exhausted and quit, and that he launched a formalization project out of frustration, to get around them. That project, Flyspeck, produced a fully machine-checked proof in 2014, sixteen years after the original announcement.
So yes, you can have certainty without a human who understands the argument. Notice the price. It took sixteen years. And it does not remove trust from the picture. It moves it. You now have to trust that the formal statement fed to the checker is the statement anyone cared about, and that the checker itself is sound. Someone human still decides that the sphere-packing question was worth sixteen years.
The part that is not in dispute
I have argued at length elsewhere, and at greater length in the book, that verification becomes the binding constraint whenever machines produce more than people can check. First Proof is the sharpest evidence for that claim I have seen, and it deserves its own post rather than a paragraph here, so I will leave it for one.
What belongs here is a different observation, and it survived every position above.
Levin said, almost in passing, that she still does not see the machine asking the questions. Strogatz agreed, and added that we will know they have arrived when one of them turns up as a guest on the show.
That is the whole thing, and it is worth stating flatly. The machine disproved the unit distance conjecture. Erdős posed it. Nobody has built a system that decides which question is worth eighty years of attention, and the mathematicians running First Proof have said in print that they do not yet know how they would even measure such a thing. You cannot benchmark taste when nobody can specify in advance what a good question looks like.
The mathematical community, in the IMU-endorsed Leiden Declaration of June 2026, has now written down formal commitments to keep that work human, retaining responsibility for correctness, insisting on attribution, and protecting the autonomy to choose which questions matter. Note that Strogatz is a signatory.
Erdős is the right person to end on, and Kipnis is right to reach for him. He published with hundreds of collaborators, and his rarest talent was not proving things. It was knowing what to ask, and knowing whom to ask it of. He had his own vocabulary for the profession. A mathematician who stopped doing mathematics had died. A mathematician who died had merely left.
The risk in front of us is not that machines will prove theorems. They will, and some of those theorems will be beautiful, and the field will be richer. The risk is that we quietly stop doing the part that was never about proving, because it is slow, unmeasurable, and impossible to put on a dashboard. Choosing the problem. Explaining the result. Deciding it mattered.
That is not only a question for mathematics. Every organization now running these systems faces a smaller version of it. The machine will hand you an answer. Somebody still has to have asked the right question, and somebody still has to be able to tell whether the answer is any good.
This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.
Comments
2 responses to “The Machine Proved It. Did It Do Mathematics?”
[…] The Machine Proved It. Did It Do Mathematics? A digression, and my favorite of the three. An AI model recently disproved a conjecture the mathematician Paul Erdős posed in 1946, reaching across the field into tools no human specialist would have thought to try. The machine produced the proof. It did not choose the question. The post works through why understanding isn’t decoration but compression, the way a bounded human mind fits something enormous into the space of a single brain, and why the usual “trust the result you can’t follow” analogy from medicine breaks down in mathematics, which has no second way to verify a claim besides the proof itself. The closing point holds up under the whole argument. Nobody has built a system that decides which question is worth eighty years of human attention. [Weigh in on LinkedIn…] […]
[…] I keep finding everywhere AI touches real work, and I wrote about its purest form in mathematics in an earlier piece. Cheap generation does not remove the bottleneck. It moves it downstream and makes it the whole […]