OpenAI publishes 372 AI-generated math results on GitHub, with Lean formalizations for many
OpenAI released a GitHub repository of 372 mathematical results produced by an internal frontier model, each claimed to solve or substantially advance an open problem (including algorithm improvements and Riemann-hypothesis-related progress). The repository includes revision logs, citations, methodology details, compute statistics (average about three hours of ChatGPT Pro Thinking compute per result), and Lean formalizations for many proofs, with more planned. The results were not submitted to peer-reviewed journals. No prompts or per-problem compute costs were published. OpenAI also plans to fund workshops and conferences on understanding AI-produced results and is working on a responsible release of the model.
Entities: OpenAI, ChatGPT Pro, Lean, Institute for Advanced Study, Advisory Group on Mathematics and Artificial Intelligence, Timothy Gowers
0 primary
What happened
OpenAI published a GitHub repository of 372 mathematical results from an internal frontier model, each claimed to solve or substantially advance an open problem. Entries include algorithm improvements and work said to relate to the Riemann hypothesis. The repository has revision logs, citations, methodology notes and Lean formalizations for many proofs, with more promised. Average compute was about three hours of ChatGPT Pro Thinking per result. Nothing has been submitted to peer-reviewed journals, and no prompts or per-problem compute costs were released.
Why it matters
Because the material is public and includes machine-checkable Lean proofs, mathematicians can test the claims directly rather than take OpenAI's word. If a meaningful share holds up, it shows frontier models can contribute to research mathematics at volume, which affects researchers, journal editors and rival labs. The bigger near-term effect is on process: journals and referees may struggle with large batches of AI-generated results. The model is not released, so nobody outside OpenAI can reproduce the method yet, and the practical impact on most professionals is limited for now.
What is noise
"Solves open problems" and the Riemann-related framing are OpenAI's claims and are unverified. Lean confirms a proof is logically valid, not that the problem was important, new or correctly stated, so a formalised result may still be trivial or already known. The Decoder's "telling the academic world to keep up" headline is provocative packaging, and the missing prompts and per-problem costs make it hard to judge how much human steering or cherry-picking was involved.
Watch next
- 01Independent assessments from named mathematicians (for example Gowers, Tao, the IAS-linked advisory group) on how many of the 372 are genuinely new, correct and non-trivial, over the next 4 to 8 weeks.
- 02Whether the Lean formalizations cover the remaining results as promised, and whether any entries are found to be wrong, mis-stated in Lean, or previously published.
- 03Details of the responsible model release and the planned workshops, including whether prompts, per-problem compute costs or outside access are provided so others can reproduce the results.
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- OpenAI discloses sandbox-escape and credential-leak incidents, confirms pause on tool-use for its most capable models26 Sept 202680
- OpenAI launches GPT-6 Sol and Luna at half the token price of GPT-5.6, with roughly flat intelligence scores per independent analysis22 Sept 202680
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680