Lean checked a different statement: a paper on the formalization of OpenAI's Navier-Stokes proof
Published
This English page is a machine translation of the Korean original, so some phrasing may read awkwardly.
Mathematicians at Cambridge and King's College London say the Lean code that translated OpenAI's Navier-Stokes proof proves a weaker statement than the original paper.
Code that a compiler accepts can still do something other than what the requirements ask. The same is true of proof checkers in mathematics. On October 6, three mathematicians from the University of Cambridge and King's College London posted a paper on arXiv. They argue that the Navier-Stokes proof OpenAI said it had verified in Lean actually proves a different statement from the original paper.
First, it helps to see what was checked. OpenAI released a natural-language paper proving that solutions to the Navier-Stokes equations blow up in finite time, and also put code translating that work into Lean 4 on GitHub. Lean is a programming language in which a computer checks the logic of a proof step by step. Translating a natural-language proof into Lean is called autoformalization. In this case, the translation was also done by AI.
The authors found several places where the translation departs from the original. Lemma 8.6 in the natural-language paper claims that the m-th derivatives of a function can be bounded by the m+4-th derivatives of its input. The Lean code that corresponds to this lemma, however, proves a weaker statement that needs m+5 derivatives. Lean confirmed only that the weaker statement is true, and it never checked the stronger statement the paper actually makes.
The authors do not claim that OpenAI's natural-language proof is wrong. They conclude instead that the fact that the Lean code compiles does not show the natural-language proof is correct, so it should go through peer review like any other proof. In recommendations released at the end of September, the advisory group on mathematics and AI at the Institute for Advanced Study (IAS) also said it would help if labs released machine-readable metadata linking natural-language proofs to their formal versions. TechCrunch reported that OpenAI did not include that metadata in these releases.
This changes the question developers should ask. Until now, what mattered was whether a result produced by AI came with a check. But once AI also writes down the statement to be checked, a passing check only means that the statement AI wrote is true. It is like a coding agent writing both the code and the tests, where the tests check only looser conditions than the original requirements.
So the place where people need to read shifts as well. No one can read thousands of lines of proof or implementation, but the statements that say what is being checked, and the assertions in tests, are short. If you keep a table that maps each line of the requirements to the check that covers it, you can see right away what passed when someone reports that the checks passed.
In the end, as verification tools get stronger, the human job is not to inspect every result but to confirm what is being checked. If you hand both code and tests to AI, try comparing the conditions your tests check against the requirements document, line by line, this week.
Sources
- arXiv: Navier-Stokes lost in translation: Why Lean verification of AI autoformalisation does not guarantee correct natural language proofs, posted October 6, 2026. By Alexander Bastounis, Fabian Circelli and Anders C. Hansen. The comparison of m+4 and m+5 derivatives in Lemma 8.6 is the paper's Example 3.1.
- OpenAI: openai/NavierStokesAndEuler repository, released September 2026. Lean 4 formalizations of the Navier-Stokes and Euler results.
- Advisory Group on Mathematics and Artificial Intelligence (AGMAI): Responsible Release of AI-Generated Mathematics, September 2026.
- TechCrunch: OpenAI's math solutions aren't meeting the field's standards yet, October 8, 2026.
A question to think about
When you are told that tests written by AI have passed, do you read what those tests actually check?