OpenAI releases 722 AI-written math papers on GitHub. Of 372 groups, 235 have Lean documents
Published
This English page is a machine translation of the Korean original, so some phrasing may read awkwardly.
OpenAI has released 722 math papers produced by an unreleased internal model on GitHub. Of 372 groups, 235 link to Lean documents that let a computer check the proofs.
In mathematics, a new proof usually goes through review by other mathematicians before the field accepts it. On October 6, U.S. time, OpenAI uploaded 722 math papers produced by an unreleased internal model to GitHub all at once. Grouping related papers together gives 372 groups.
OpenAI said it asked the model to solve about 4,000 problems, then selected only the results it considered significant enough. On average, each result used compute equivalent to about three hours of thinking by ChatGPT Pro. According to Scientific American, OpenAI said almost all the results came from giving a single prompt to a single agent.
Lean provides a way to check the results. It is a programming language that lets a computer check the logic of a proof step by step. Once a proof is translated into Lean, its logic can be checked without a person having to read it. Counting the entries in OpenAI's paper list shows that 235 of the 372 groups have Lean documents, while 137 do not. OpenAI also said results that have not been translated into Lean may contain problems.
There are exceptions. OpenAI said it used a different process for its work on zero-free regions of the Riemann zeta function and a proof of a special case of the Hodge conjecture. Both topics are connected to famously difficult problems. The zeta function write-up was edited by humans to make it easier to read.
Others cannot reproduce the results using the same process. OpenAI has released neither the model nor the prompts. The mathematics and AI advisory group at the Institute for Advanced Study (IAS), which OpenAI consulted, recommended releasing the model, the exact prompts, and the compute time for each problem. According to Scientific American, OpenAI disclosed only the average compute time and a few statistics. MIT's Andrew Sutherland said the claim that a single agent solved the problems in one attempt should be treated as unverified until the model is released and the results are reproduced.
The way we assess this news needs to change, too. Until now, the fact that AI had solved a difficult problem was itself news. But when 722 papers arrive at once, the first question is who checked the claims and what they verified, rather than what the AI claims to have solved.
Developers will recognize the situation. When an AI-generated pull request comes with tests and type checks, a machine can catch problems first. Without those checks, someone has to read through it and assess it from scratch. When assigning work to AI, we need to consider whether we can also build checks that automatically verify the result.
Lean does not guarantee everything, though. It checks whether a proof correctly establishes the statement it sets out to prove. A person still needs to check whether that statement faithfully represents the original problem. Passing tests does not necessarily mean the requirements have been met.
So whether we can trust AI's results will depend more on what has been verified than on how many results it produces. When assigning work to AI, the first step is to decide what a machine will check and what a person will need to read, before the results arrive.
Sources
- OpenAI: Sharing AI progress in mathematics, October 6, 2026. Automated access to this page was blocked, so the announcement was cross-checked against sources 3 and 2, which carry the same announcement.
- OpenAI: openai/math repository, October 6, 2026. README, paper list (CONTENTS.md), and Lean formalization list (lean/formalization.yaml). The count of 235 groups with Lean documents comes from directly counting the Lean links in the paper list.
- OpenAI Community: First look at mathematics manuscripts from an internal frontier model at OpenAI, October 2026.
- Scientific American: OpenAI unleashes hundreds more math results upon a field already in shock, October 6, 2026. By Joseph Howlett.
A question to think about
Of the tasks we give AI in our services, how many have results that a machine can check for correctness?