Introduction
If you don’t know what this blogpost is about, my congratulations to you, and please let me know if there is some additional place under that rock.
On 8 September 2026, artificial intelligence company OpenAI announced a proof of a breakdown of Navier–Stokes solutions in three-dimensional Euclidean space, developed by its researchers using thousands of agents running an internal frontier model, along with a formalization in the Lean proof assistant. The claim has not been verified by external mathematicians or the Clay Mathematics Institute, while OpenAI stated it would not claim the Millennium Prize. The announcement was accompanied by a priority dispute with Levent Alpöge (employed at rival AI company Anthropic) and Tristan Buckmaster, who had derived a set of closely related results on the Euler equations. The method used to generate the claimed solution built upon a method developed by Diego Cordoba and Luis Martinez Zoroa in 2023 to prove blowup phenomena in related fluid equations.
Broken record alert: this is only one of an infinite number of explanations of what just happened. You can spend the better part of a day/week/month looking up reports of what happened, what Navier-Stokes (the equation) is, why it matters, and what this latest proof means, and you probably should spend a little bit of time chatting about this with your LLM of choice.
But Who Did It?
The interesting thing about this proof isn’t just the fact that it was discovered. The who and the how of it is what is truly fascinating, and going through this aspect is what today’s blogpost is about.
First, the controversy:
The night before OpenAI announced the proof, Buckmaster posted on social media that he and Alpöge, who is a mathematician and an employee of Anthropic, had blown up the Euler equations, widely seen by the field as a step toward solving the Navier-Stokes problem. In a statement released on Monday, Buckmaster alleged that OpenAI had found out about the pair’s progress in the last week and had adopted the same method as Buckmaster and Alpöge had been using to solve the full problem. In a press briefing on Tuesday, OpenAI mathematician Sébastien Bubeck denied the rumors about the proof’s origins. Bubeck said OpenAI’s internal model had independently solved the Euler problem by totally different means than those employed by Buckmaster and Alpöge. OpenAI’s solution to the full Navier-Stokes problem, however, did follow a similar method as that used by the two mathematicians. And the proof was developed over the weekend, according to Bubeck—that is, after the time that Buckmaster claims the news of his and Alpöge’s result had reached OpenAI. Bubeck emphatically denied that the two mathematicians’ work had influenced OpenAI. “We did not use their prompt or proofs to prompt our models,” he said.
Again, there’s a lot more where that came from. If you choose to drill down into the details, you can spend an additional day/week/month on this aspect of the story. It is, after all, the age of abundance!
So who did solve it? Did Buckmaster and Alpöge (team Anthropic) solve it? Or did Sébastien Bubeck (team OpenAI) solve it?
Or did Claude and ChatGPT solve it?
Or - and this is my preferred way to phrase it - did the humans figure out that this particular problem was worth solving, and then the models went and solved it?
In other words, how much credit should the models get for coming up with the solution? When I solve 432*978 using my calculator, who has solved the problem? Me, or the calculator? I think it is fair to say that I decided that solving this particular problem was important, and the calculator did the calculation.
But the Navier-Stokes thing is more complicated than that. And that brings us to the second question:
How Did They Do It?
Go back to my toy example (432*978). Seriously now: who solved it? Technically speaking, of course the calculator solved it. But again, most humans will accept that the facts that:
a human decided that these two numbers were the important and relevant ones, and
that a human decided that the mathematical operation to be performed upon them was multiplication
…means that the human “was in charge”.
So, with Navier Stokes, who was in charge? Lengthy excerpt follows:
On Tuesday, September 1, we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors and by the step change in performance of our internal model, we launched an effort to evaluate it on all open Millennium Prize problems and a few other high-impact problems.
We used a system of coordinating agents powered by our internal model. The agents had access to tools such as the ability to read from a cached version of the internet and the ability to run code. Agents were subdivided into groups with the ability to communicate within the group. The groups varied in size, and the group that produced the Navier–Stokes resolution involved on the order of 10,000 concurrent agents. At all times we maintained the same strict safeguards that we apply to all our frontier model evaluations, including monitoring and isolation.
For each problem, we prompted different groups of agents with different variants of the problem statement, covering all variants of the problem. For the Navier–Stokes problem, we suggested versions “A” and “B” (particular forms of the Navier–Stokes problem which would result in a proof) and versions “C” and “D” (which would result in a disproof) to separate groups of agents.
In addition to the full Millennium Prize problems, we asked our multiagent system to try a set of “easier” problems. One of these problems was a similar blowup question for the limit of the Navier–Stokes problem with the viscosity term removed. This is known as the regularity problem for the Euler equations, and our agents surprised us by resolving this question. The specific variant of the question that they resolved was the unforced version, where no external force is applied to the fluid. Nearly 100 agents worked together for approximately 50 hours to produce our Euler regularity disproof.1
Once we saw the Euler solution, we thought that Navier–Stokes was the most promising problem to work on. Thus, we decided to devote our resources to Navier–Stokes. To do so, we shifted agents away from the other Millennium Problems and prompted these agents with the Euler resolution. When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model.
We encouraged different groups of agents to explore a diversity of approaches. After some time, we cross-pollinated the agent groups by using Codex to consolidate the most useful insights from each agent group. These follow-up prompts drew on the agents’ own intermediate results. The group that found the solution to Navier–Stokes was guided in such a way.
The agents arrived at their resolution on Saturday, September 5, about 88 hours after the first agents were launched. Lean formalization and verification took an additional 17 hours via GPT‑6 Astra.
Across all attempted problems, the agents sent 4.9 million messages and used about 300 billion output tokens. In the process of resolving the Navier–Stokes problem, the agents sent 2.7 million messages and used approximately 130 billion output tokens.
Phew, sorry about that. But hopefully, you found it as interesting as I did. And if you did, here’s a thought experiment for you:
What if I point my Codex/Claude Code at this article, an excerpt from which I have shared with you above
I say to my Codex or Claude Code on Monday morning: “keep a track of the most exciting, passionate discussions about unsolved, important and application oriented problems in mathematics, and do this by figuring out what sources will be best for such an exercise. Once a week, let us review your sources, and you and I will improve them over time. Each week, figure out a list of the ten most exciting problems (howsoever you define them) to work on. Of these ten, select the three that are in your opinion most tractable. Of these, start working on the first one, and if you haven’t solved it by Tuesday evening, move to the second. If you haven’t solved it by Thursday morning, move to the third, and if you haven’t solved that one by Sunday evening, report back to me, and we will rejig the process and start afresh next Monday. If you do solve any one, move on to the next one on your list”.
When my Codex/Claude Code have solved a problem, does it deserve a prize, or do I?
This may seem like a slight to the humans involved in the current controversy, but that is not at all my intention. I am saying that in the very near future, this is likely to be reality, and a quotidian lived experience. That prompt I’ve written above may (will!) actually work. Does that mean I get the credit? Should I?
What do - and I mean this question entirely seriously - credit and status even mean in the age of AI?
Wallace and Darwin, But Fast-Forwarded
Did Darwin come up with the theory of natural selection, or did Wallace? That remains a controversy to this day (that page I’ve linked to makes for fascinating reading!), although it is fair to say that both Darwin and Wallace handled the issue with more grace than has been on display over the past couple of days.
But that abundance of grace, it is fair to say, is also a function of the time available back then. The controversy, such as it was, (and a litany of other tragedies besides) unfolded over the course of months, not weeks, and certainly not days. And the issue at heart was simply about who discovered the theory first, not about a whole host of egos, balance-sheets, IPOs and model capabilities. Natural selection, you could say, simply hadn’t had time to work on the evolution of controversies.
But the larger point I want to make in this section is that it is not just the fact that this particular controversy (and therefore the underlying discovery) is fast-forwarded - it is that everything is fast-forwarded.
We live in an age where calling ten thousand agents to help you is trivial, but verifying the work that those ten thousand agents do can take years. In my current day job, I automate legal processes using AI, and we face the same problem (though not, thank god, the same level of mathematical abstruseness): automating the process is the relatively easy bit. Getting a trained human being to verify the quality of the automated output is the hard part.
And so in a weird way, the fight over who-invented-it-first is all-too-understandable. There are likely to be many more discoveries in the days/weeks/months to come, but it will be increasingly difficult for humans to credibly take credit for ‘em.
A Most Wondrous Mind, Unable to Wonder
Many years ago, I read a nice little book in the British Council Library. It was called The Age of Wonder, and it was written by Richard Holmes.
“The Age of Wonder is a colorful and utterly absorbing history of the men and women whose discoveries and inventions at the end of the eighteenth century gave birth to the Romantic Age of Science”, says the blurb on the Amazon page.
We live today in an age where the discoveries and the inventions that are being made will be infinitely more colorful and absorbing. But the inventors and the discoverers themselves, alas, will not be quite as colorful, nor as utterly absorbing.
And so, in the spirit of celebrating that which is scarce, I say we should be forcibly grateful for the controversy surrounding who discovered the forced Navier Stokes thingummy.
They won’t be around for much longer, you see, these controversies about discoveries made by humans. So please, bring it on, and may the matter never be resolved to everyone’s satisfaction.


