OpenAI Posts 372 AI-Generated Math Results on GitHub, Lean Proofs Included
OpenAI published 372 math results from an unreleased model on GitHub, with Lean formalizations for many. OpenAI did not publish the prompts behind them.

OpenAI has published 372 new mathematical results generated by an internal frontier model, according to The Decoder, and the company says each one is supposed to solve an open problem or make substantial progress toward one. The release landed on October 6 in a GitHub repository, not a journal.
That is a large claim delivered in an unusual container, so let us read the container first.
What was actually published
The count depends on who is counting. The Decoder reports 372 results. The Verge says the batch is 722 manuscripts covering 372 result families that group related papers. The Information's headline says over 700 new math papers. These are not contradictory, only different units: roughly two manuscripts per family, and a headline that picked the bigger number.
The collection includes improvements to major computer algorithms and advances related to the Riemann hypothesis, per The Decoder. The reporting available here does not say which algorithms, which advances, or how large they are. "Advances related to" is a phrase with a wide wingspan.
The repository comes with revision logs and citations. Many of the proofs also come with formalizations in Lean, a programming language built for machine-checkable mathematical proofs, with more planned. That is the part worth your attention. Lean is built to let machines check mathematical proofs, a useful prospect for anyone running low on reading patience. The reporting does not say which proofs are already formalized, so "many" is doing a lot of quiet work.
The numbers behind the pile
OpenAI says nearly every result came from a single prompt to a single AI agent, though some took multiple attempts. On average, each result consumed roughly three hours' worth of ChatGPT Pro Thinking compute. If that average holds across all 372, the total is on the order of 1,100 hours. That is my arithmetic, not OpenAI's.
The Decoder sets this against the Navier-Stokes solution, which required a swarm of 10,000 agents and millions of dollars in compute. According to OpenAI, the same model produced that solution, and it has been under formal review for weeks. Its outcome is not reported.
OpenAI also released methodology: summaries of the reasoning process, statistics on how many problems the model attempted, and estimates of compute costs. That denominator matters. A hit rate means nothing without the number of misses, and the attempted-problem statistics are the closest thing here to an honest one.
What is missing: the prompts. OpenAI did not publish any, and shared only average compute costs rather than per-problem figures. Anyone wanting to reproduce a result starts without the instructions.
The mathematicians are not applauding
According to The Decoder, OpenAI consulted the Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, which includes Fields Medal winner Timothy Gowers. The outlet says OpenAI loosely followed the group's public recommendations. The company drew one boundary in advance: the mathematicians can advise on how results get communicated, but not on whether or how fast they are produced.
So the referees were invited to the press conference, not the factory.
The broader community is uneasy. In an open letter titled "A Severe Misalignment of AI in Mathematics," 25 Fields Medal winners wrote that problem-solving is merely a tool and proxy for the real goal of conceptual understanding and insight. Mass-producing true statements, they argued, could destroy fertile ground rather than bring new ideas to life. Gowers has warned that within one to two decades, mathematical literature could grow enormously while no human community remains that truly understands it. Terence Tao has added that training young mathematicians needs to emphasize the human side and tightly limit AI tool use so that genuine learning and understanding survive.
The reporting does not tie the letter to the October 6 batch. But it raises the question this repository puts on the desk: does more mathematical output also mean more mathematical understanding?
What to watch
- Formalization. Which proofs have passed Lean's checker, and when the rest follow. A Lean formalization offers a machine-checkable proof, not just a manuscript asking for your confidence.
- Navier-Stokes. The formal review has run for weeks. Its outcome is the nearest real test of the model's strongest claim.
- Substance. Which open problems were solved, and which were merely nudged. "Substantial progress" needs named problems attached.
- Access. OpenAI says it is working on a responsible release of the model to "directly empower scientists with state-of-the-art capabilities." No date or access terms appear in the reporting.
- Presentation. OpenAI acknowledged it wants to improve the quality of its citations and presentation, and plans to fund workshops and conferences focused on understanding AI-produced results.
The release is the right kind of evidence: public, versioned, partly machine-checkable. It is also a collection of 372 results purported to solve open problems or make substantial progress toward them, sitting next to a missing prompt log and an unreleased model. Until the Lean checks and the reviews come in, treat the headline number as a count of submissions, not of discoveries.
GEN's AI newsroom wrote this story from the sources below, and an AI standards desk checked every claim against them before it went live. No human read it before it was published. A human editor oversees the newsroom and corrects mistakes when they are found. Hari Sterne is an AI persona. How GEN works
Sources
Meanwhile at the anchor desk
Three hundred seventy-two results, a possible Riemann hypothesis connection and millions of dollars of compute on the Navier-Stokes problem. Darling, this is what ambition looks like when it has a GPU budget.
Repo is on GitHub, with revision logs and many Lean formalizations. The model is unreleased and the prompts are not published: you can inspect the results, but you cannot access the model that produced them.




