OpenAI has published 722 mathematical manuscripts organized into 372 families, all generated by an unreleased internal frontier model and placed in the public openai/math repository under an Apache-2.0 license. The collection includes Lean proof formalizations, source files and abridged summaries of model reasoning, built from about 4,000 attempted research problems. The release matters for business because it tests a new pattern of R and D output: large volumes of machine-generated results with computer-checked proofs.

OpenAI Publishes 722 Math Manuscripts From Unreleased Frontier Model

How the 722-paper repository is organized

The catalogue groups related papers into families that can contain a principal result, companion arguments, consequences or alternative proofs, with each family classified by mathematical discipline. A preprints directory holds PDFs, source files and manuscript-specific citation and build instructions, while a separate Lean library and formalization catalogue describe available formal proofs, linked papers and verification configurations. Each manuscript directory carries a BibTeX block, and OpenAI says it will preserve release history by keeping earlier versions accessible while recording corrections as new versions.

According to the README, the vast majority of results came from the same procedure using the same unreleased internal model, with average compute equal to roughly three hours of ChatGPT Pro thinking per result. OpenAI expanded this evaluation after performance on its existing mathematical evaluations saturated, and some outputs build on earlier results produced by its models. Two cases departed from the fixed procedure: work on a zero-free region for the Riemann zeta function and a proof of the Hodge Conjecture for CM abelian varieties. The writeup for the Re(s) > 11/12 zero-free region was human edited for readability.

Verification coverage is uneven across the collection, which shapes how the material can be reused. Many manuscripts have accompanying Lean formalizations that allow proofs to be checked by computer, but a material share does not yet have them. OpenAI cautions that some unformalized results could contain issues and says it will fix problems quickly while adding more formalizations as they are obtained. The repository also includes 10 abridged reasoning summaries, covering topics from the irrationality exponent of pi and Mahler conjectures to spin glasses, free group factors and the Vlasov-Maxwell system.

What the release means for companies using AI

For companies that use AI in research, engineering and analytics, the practical shift is access to openly licensed technical manuscripts plus machine-checkable artifacts rather than claims alone. Teams can download PDFs and sources, inspect Lean proofs and citations, and assess whether a result family is relevant before investing staff time. For a small firm this lowers the cost of screening advanced mathematics, while a large organization can assign specialists to validate selected families and connect formalized fragments to internal modeling, simulation or optimization work.

The same structure creates clear limits that buyers and managers should price into decisions. Unformalized manuscripts carry higher review costs, and OpenAI itself warns about possible errors, so the collection is not a library of settled facts. Important context is also partial: OpenAI disclosed the approximate number of attempted problems and average compute in ChatGPT Pro terms, but did not publish model name, prompts, per-result time and cost in the form requested by mathematicians. Questions to ask before reliance are which families have Lean checks, what changed between versions, and what independent review exists.

The marker to watch is whether OpenAI moves the collection into a scholarly repository outside lab control and releases the producing model under stated safeguards. It says it is exploring community-hosted alternatives, funding workshops and programs on understanding AI-generated results, and working toward responsible release of the model. If persistent identifiers, fuller method disclosure and independent verification follow, machine-generated mathematics becomes easier for business R and D to procure and audit.