Category: Digressions

Pieces outside the enterprise frame that carry the same underlying question.

  • Could a Machine Have Had Darwin’s Idea?

    In November 2020 I started a thread on X. I have barely used the platform these past couple of years, and I only came back to this thread recently, by accident, when some of my own old posts surfaced in front of me. Reading them in one sitting was strange, because a thread I had added to piecemeal over five years turned out to have been about one question the whole time.

    It opened with Darwin. Someone I followed had posted, amazed, about what evolution manages to build, and I wrote that what amazed me was something else: the intellectual leap required to infer the theory in the first place, from nothing but empirical observation and years of focused study. I said I was not sure I fully appreciated Darwin’s intellect and dedication. Then, in the next post, I said the thing the whole thread has been chasing ever since. That kind of leap, I wrote, is the kind of intelligence AI should be aiming for, not recognizing cat faces or copying tasks people already do well.

    That was the bar I set in 2020, and it was a high one. Darwin spent decades buried in finches and barnacles and pigeon breeders and fossil beds, a mountain of unconnected observation, and then made a leap: one idea, natural selection, that reorganized all of it at once. That inductive move, from a heap of messy particulars to the principle that explains them, struck me then as the most formidable act an intellect can perform, and the part of real science I was least sure a machine could touch. So I kept a list, adding to it whenever I came across a case of AI looking like it was helping advance science rather than just crunch it, to watch whether anything ever cleared the bar.

    Reading the thread back now, it traces an arc I did not plan, and it worried at the right problem from the start. Within days, in November 2020, I posted a New York Times piece on Max Tegmark’s group, whose neural network had recovered a hundred physics equations from raw data, and I pulled out the catch that the physicists themselves named. Tegmark was candid that the machine could retrieve the formulas but not yet the deep principles beneath them, the quantum uncertainty or the relativity that would explain why the formula holds. And Jesse Thaler, the MIT physicist directing the new AI-and-physics institute, put his finger on why. AI wins at games because a game has a well-defined notion of success. “If we could define what success means for physical laws,” he said, “that would be an incredible breakthrough.” Proposing the theory, in other words, was the hard part, and it was hard precisely because you could not say in advance what would count as getting it right.

    The entries that moved me most in those early days were not about AI at all. They were about humans doing the thing I wanted to see a machine do. Around the same time I was reading about Tibor Gánti, the Hungarian biologist who deduced from first principles what the simplest possible living thing must be: a metabolism, a way to store information, and a membrane, three systems that have to be coupled or the organism dies. I wrote at the time that this was the kind of inductive thinking we should be teaching children. Later I followed a related idea, assembly theory, which proposes a single elegant handle on complexity: count the minimum number of steps needed to build a molecule, and use that number to test for the presence of life on other worlds. A theory reduced to one measurable quantity, aimed at one of the hardest questions there is. It may not survive contact with the evidence; I later noted that a study found some minerals scoring above the threshold the theory sets for life. But right or wrong, it was the move I admired, the leap from observation to a principle sharp enough to be tested.

    Then the machine examples accumulated. In December 2021, a Nature paper where neural networks guided the intuition of mathematicians toward new conjectures in knot theory, the machine surfacing the pattern and the human still making the leap. In March 2022, a system that rediscovered Newton’s law of gravitation from the motion of the planets and wrote it back out as a symbolic equation. Then more symbolic regression, pulling laws from data. Then, in February 2025, the entries that made me sit up: Google’s AI co-scientist, generating and ranking novel research hypotheses on its own, and Evo-2, a model that does not just read genomes but writes them. And I drew a line I did not fully understand at the time. I noted that the most exciting work was different from the systems that merely “generate plausible hypotheses from an extremely large space of possibilities.”

    Five years of watching, and that distinction turned out to be the whole story.

    The step that was supposed to have no method

    For a century, the standard account of science drew a hard line between two acts.

    One is coming up with the idea. The other is checking whether the idea is true. Karl Popper named these the context of discovery and the context of justification, and he was blunt about which one belonged to philosophy. Testing a hypothesis has a logic. Having one does not. In his words, the act of conceiving a theory “neither calls for logical analysis nor is susceptible of it.” The initial leap from a pile of observations to this might be why was, he thought, a matter for psychology, not method. A hunch. The unteachable part.

    That is where the whole romance of science lived, and Darwin’s leap is its patron saint. Kekulé dreaming the benzene ring as a snake biting its tail. Fleming noticing the one culture plate that had gone wrong in an interesting way. Darwin holding twenty years of specimens in his head until they resolved into a single idea. We told these stories because the leap seemed to come from nowhere, and coming from nowhere was the point. You could train someone to run an experiment. You could not train the hunch.

    The systems in my thread are automating the hunch.

    Not perfectly, and not everywhere. But the co-scientist does not summarize the literature and hand you a reading list. It proposes mechanisms no one has written down, argues them against itself, and ranks the survivors. Evo-2 does not retrieve a gene; it composes one. Whatever you want to call that, it is happening on the discovery side of Popper’s line, in the territory he declared off-limits to method. The unteachable step is being done by a machine that was, in fact, taught.

    What actually got cheap

    Here is where my thread stops being a highlight reel and starts being an argument, because the interesting question is not whether machines can generate hypotheses. They plainly can. The question is what that does to the rest of science.

    The mathematician Noah Giansiracusa has a name for the pattern, which I take up at more length in my book: carpet bombing. When generation gets cheap, you stop being clever about producing candidates and start producing all of them, then sort. It is how AI does mathematics, throwing enormous numbers of attempts at a problem where checking each one is fast. Hypothesis generation is now carpet bombing pointed at nature. The co-scientist can produce more plausible, well-argued, literature-grounded hypotheses in an afternoon than a lab could dream up in a year.

    And that is exactly where the trouble starts, because a hypothesis is not a proof. It cannot be checked in an afternoon. It has to be checked against the world, and the world runs on its own clock.

    When generation was expensive, the scarce, precious act was having the good idea, and verification, while never easy, was not the binding constraint. A scientist had three hypotheses worth testing and a career to test them in. Reverse that. Now the machine hands you three hundred plausible hypotheses, and the binding constraint is the wet lab, the clinical trial, the telescope time, the years. Generation raced ahead. Verification did not move at all, because verification in science is not a faster model. It is reality, taking as long as reality takes.

    This is the same shape I keep finding everywhere AI touches real work, and I wrote about its purest form in mathematics in an earlier piece. Cheap generation does not remove the bottleneck. It moves it downstream and makes it the whole game.

    The field is already learning this the hard way

    You do not have to take the argument on faith, because the correction is already arriving, and it is arriving in the most useful form: from the people who built the tools.

    Google’s co-scientist reached Nature in 2026, with real wet-lab validation in a handful of biomedical cases. Impressive, and I do not want to wave it away. But when an independent researcher carefully re-implemented the system and ran it hard, the finding was sobering. The pipeline reliably produces hypotheses. Whether it actually improves on the underlying model’s raw guesses was not reproducible from one run to the next, and across dozens of attempts on one disease, not a single one of its generated hypotheses matched the paper’s own headline discoveries. The machine is a fountain of plausible ideas. Plausible is not the same as true, and telling them apart is still the expensive part.

    There is a sharper cautionary tale two years older. In 2023 Google reported that around forty new materials had been discovered and synthesized with the help of one of its AI systems. It was held up as a landmark. Then outside chemists went through the results, and an independent analysis concluded that not one of them was actually a net-new material. The generator worked. The verification, done properly and after the fact by humans, is what separated the discovery from the illusion of one. Every honest account of these systems now carries the same caveat, in the developers’ own words: careful experimental validation, peer review, and independent scrutiny are what turn a generated candidate into knowledge.

    That caveat is not a footnote. It is the job.

    Which part was ever the science

    So let me go back to my thread, and to the line I drew without fully appreciating it.

    I think I was reaching for this, and the clue was in Thaler’s line about defining success. The systems I found most exciting were not the ones that produced the most hypotheses. They were the ones tied to a way of checking, the model that rediscovered gravity and could be tested against known physics, the mathematical work where a conjecture could be pursued to a proof. The ones that unsettled me were the pure generators, magnificent at producing possibilities and silent on which ones were real. The 2020 worry and the 2025 worry are the same worry. A physical law you cannot define success for, and a hypothesis you cannot yet verify, are the same problem.

    Popper drew his line to protect justification. He wanted to say that the logic of science lived in the testing, and that the having-of-ideas, however romantic, was not where rigor lived. The machines have now inverted his world in the most ironic way possible. They have automated the part he thought had no method, the hunch, and in doing so they have made the part he cared about, the checking, more valuable than it has ever been. When hunches were scarce, verification could feel like bookkeeping. Now that hunches are infinite and nearly free, verification is the only thing standing between a lab and a year spent chasing a beautifully argued hypothesis that was never going to be true.

    There is a clue to this in something Andrew Wiles once said, that it is bad to have too good a memory if you want to be a mathematician. It sounds backward until you see what he means. The mathematical gift was never recall. It was compression, the knack for throwing away almost everything and keeping the one idea that organizes the rest, which is roughly how Jürgen Schmidhuber defines insight: a better, shorter way to predict what you have seen. A machine with perfect memory and unlimited generation has exactly the strength Wiles warns against and not yet the one he prizes. It can hold everything. It cannot yet tell what to forget.

    Which returns me to the question I started the thread to answer. Could a machine ever do what Darwin did?

    I think the honest answer is now a qualified yes, and it is qualified in a way I did not expect. A system can hold a mountain of observation and propose the organizing idea. It can make the inductive leap I was so sure was ours alone. But watching it happen, I realized I had misjudged where Darwin’s genius actually sat. The leap to natural selection was extraordinary, but the leap was not the science. The science was in the twenty years, in Darwin knowing which single idea out of the dozens he entertained was worth a life’s defense, and in his relentless testing of it against every objection he could invent. The hunch was the cheaper half but we just could not see that while hunches were rare.

    So the machines got the hunches, and they are welcome to them. What Darwin had that they still do not is the judgment to know which hunch was worth everything, and the patience to spend years finding out if he was wrong. That part did not get automated. It got scarcer, and more valuable, than it was when the ideas were hard to come by.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.

  • The Machine Proved It. Did It Do Mathematics?

    A digression. This blog is usually about enterprises. This one is about mathematics, and about a question I have been fascinated with.

    In May, an OpenAI model disproved a conjecture that Paul Erdős posed in 1946.

    The problem is easy to picture, which is part of its charm. Scatter some dots on a page. Count the pairs that sit exactly one inch apart. As you add more dots, how fast can that count grow? Erdős guessed there was a ceiling, and that a grid-like arrangement came close to it. For decades most mathematicians thought he had it right. The model found arrangements that beat the ceiling, and kept beating it as the number of dots grew without limit.

    The surprising part was the route. The unit distance problem belongs to discrete geometry. The solution came through algebraic number theory, which studies something else entirely. Writing in the Wall Street Journal last month, the statistician Daniel Kipnis makes an observation about this that I keep returning to. Cross-disciplinary borrowing is not new in mathematics. Descartes did it in the seventeenth century. What has changed is scale. A mathematician can spend an entire career in discrete geometry and never acquire the tools of algebraic number theory, because a human career is short and those tools take years. A machine has no such constraint. Its reach across the field is bounded by the cost of computation and nothing else.

    So the machine went somewhere no specialist would have thought to look, and came back with a counterexample.

    Then Kipnis asks the question that makes this interesting. If nobody understands a proof, is it a proof at all?

    Mathematics as a social act

    He gets his answer from Reuben Hersh, who spent a career arguing that mathematics is a social phenomenon rather than a collection of eternal truths sitting somewhere waiting to be found. On that view, a mathematical fact does not become part of mathematics by being true. It becomes part of mathematics by being discovered, explained, and absorbed into what the community understands. Progress is a form of communication. A proof explained badly does no more good than a proof that is wrong.

    Kipnis notes that OpenAI seems to have understood this instinctively. It did not publish the output and walk away. It worked with prominent mathematicians who verified the argument and wrote a companion paper making it intelligible to the field. Without that second step, the result would have occupied a strange position. It might have been true. But if no human could confirm it or follow it, what would its truth consist of? On Hersh’s account it would not yet be a mathematical result.

    This is a strong claim and I find it persuasive. It also has a problem, which two other people, in another discussion, illuminate.

    Is understanding a crutch?

    On Quanta’s podcast The Joy of Why, Steven Strogatz recently interviewed Lauren Williams, the Harvard mathematician who helped start the First Proof project. Afterward Strogatz and his co-host Janna Levin, an astrophysicist, kept talking, and the conversation turned to something more unsettling than job displacement.

    Strogatz asked whether beauty will still guide mathematics once machines are doing it alongside us. Beauty in the working sense, meaning the aesthetic pull that tells a mathematician which question is worth asking and whether an argument is on the right track. Earlier in the discussion, Strogatz and Williams had discussed about her philosophy on beauty: “If you ask a question and the answer is not beautiful, that means you asked the wrong question.”

    Levin’s answer is the best thing I have read on this in months. One of the things beauty does, she said, is make a complicated subject comprehensible. Then she gave the reason she needs that: “I don’t have infinite compute.”

    Elegance is not decoration. It is compression. Understanding is the technique a bounded mind uses to fit something enormous into a space the size of a human head. We prize proofs that are short, surprising, and clean because we cannot hold the long ugly ones. As Schmidhuber, a leading AI scientist, explains, a computationally limited observer finds something simpler and more beautiful once she learns to predict and compress the data in a better way. Herbert Simon spent a career making a version of this argument about organizations, which exist in part because no individual can hold the whole problem, so the problem gets cut into pieces a person can carry. Mathematical understanding looks like the same adaptation, running on the same constraint.

    Which invites the obvious follow-up, and Strogatz asked it. If understanding is a workaround for our limits, is it overrated? He suggested we might be confusing means with ends. If the goal is true theorems and reliable prediction, comprehension is the ladder, and once you are up you can kick it away. He offered a medical analogy. If a therapy saves a life, you may take it without understanding why it works.

    He also gave the other side its due, which is that some people regard science without understanding as a diminished thing, and he said he could see both positions.

    Where the analogy breaks

    But notice what the medical case is quietly relying on.

    You can accept a treatment you do not understand because you have another way of knowing it works. The trial. The outcome is observable, the effect is measurable, and the verification runs on a completely separate track from the explanation. Understanding is genuinely optional there, because something else is doing the job that understanding would otherwise do.

    Mathematics has no second track. There is no experiment that shows a theorem is true. You cannot run a trial on a conjecture. The only instrument the field has ever had for establishing that a statement holds is a proof, and a proof is a piece of writing addressed to another mind. In mathematics, verification and explanation are not two activities that happen to co-occur. They are the same act.

    That is why Hersh’s position is stronger than it first appears, and why “understanding may be overrated” does not transfer cleanly from medicine to mathematics. Give up on understanding a proof and you have not traded comprehension for reliability. You have given up your only method of knowing.

    The escape hatch, and what it costs

    There is one way out, and it is real. Machines can check proofs.

    This is not new and the mathematics community has been living with the discomfort for fifty years. The four color theorem fell in 1976 to an argument that included computer case-checking no human could reproduce by hand, and mathematicians argued about whether that counted.

    The sharper case is Thomas Hales. In 1998 he announced a proof of the Kepler conjecture, about the densest way to stack spheres. The Annals of Mathematics assigned twelve referees. They worked for four years. They concluded they were ninety-nine percent certain the proof was correct, and admitted they could not independently verify the thousands of lines of computer code it rested on. Full publication came nearly eight years after submission. In a retrospective written years later, Hales says plainly that the review dragged on until the referees became exhausted and quit, and that he launched a formalization project out of frustration, to get around them. That project, Flyspeck, produced a fully machine-checked proof in 2014, sixteen years after the original announcement.

    So yes, you can have certainty without a human who understands the argument. Notice the price. It took sixteen years. And it does not remove trust from the picture. It moves it. You now have to trust that the formal statement fed to the checker is the statement anyone cared about, and that the checker itself is sound. Someone human still decides that the sphere-packing question was worth sixteen years.

    The part that is not in dispute

    I have argued at length elsewhere, and at greater length in the book, that verification becomes the binding constraint whenever machines produce more than people can check. First Proof is the sharpest evidence for that claim I have seen, and it deserves its own post rather than a paragraph here, so I will leave it for one.

    What belongs here is a different observation, and it survived every position above.

    Levin said, almost in passing, that she still does not see the machine asking the questions. Strogatz agreed, and added that we will know they have arrived when one of them turns up as a guest on the show.

    That is the whole thing, and it is worth stating flatly. The machine disproved the unit distance conjecture. Erdős posed it. Nobody has built a system that decides which question is worth eighty years of attention, and the mathematicians running First Proof have said in print that they do not yet know how they would even measure such a thing. You cannot benchmark taste when nobody can specify in advance what a good question looks like.

    The mathematical community, in the IMU-endorsed Leiden Declaration of June 2026, has now written down formal commitments to keep that work human, retaining responsibility for correctness, insisting on attribution, and protecting the autonomy to choose which questions matter. Note that Strogatz is a signatory.

    Erdős is the right person to end on, and Kipnis is right to reach for him. He published with hundreds of collaborators, and his rarest talent was not proving things. It was knowing what to ask, and knowing whom to ask it of. He had his own vocabulary for the profession. A mathematician who stopped doing mathematics had died. A mathematician who died had merely left.

    The risk in front of us is not that machines will prove theorems. They will, and some of those theorems will be beautiful, and the field will be richer. The risk is that we quietly stop doing the part that was never about proving, because it is slow, unmeasurable, and impossible to put on a dashboard. Choosing the problem. Explaining the result. Deciding it mattered.

    That is not only a question for mathematics. Every organization now running these systems faces a smaller version of it. The machine will hand you an answer. Somebody still has to have asked the right question, and somebody still has to be able to tell whether the answer is any good.

    This argument runs through my book, Coherence, arriving this Fall. If you want to follow the thinking as it develops, join the list at coherise.com. The one-page decision tool from the book is the first thing I send.