AINews

OpenAI’s mathematics release faces scrutiny after three withdrawals

OpenAI withdrew three manuscripts and revised others after releasing hundreds of mathematical claims. An independent advisory group says publication starts the process of assessment.

Fuld Hall on the campus of the Institute for Advanced Study in Princeton, New Jersey.
File photograph of Fuld Hall at the Institute for Advanced Study in Princeton, New Jersey, taken in July 2023. Zeete (resized and converted to WebP). CC BY-SA 4.0.
LinkedInPostEmail
Save for later

OpenAI withdrew three mathematical manuscripts from its public GitHub collection on October 7 after identifying a sign error that also undermined dependent work. The corrections followed its October 6 release of AI-generated mathematical claims, underscoring the assessment still needed before researchers can rely on the results.

In an October 9 analysis for The Conversation, republished by Phys.org, Melissa Lee reports that the initial release comprised 722 papers relating to 372 open problems across algebra, geometry and theoretical computer science. Those figures describe the scope of the release, not independently confirmed solutions to 372 problems.

Lee reports claims of advances concerning the Riemann hypothesis and the Birch–Swinnerton-Dyer conjecture, two prominent mathematical problems. Her analysis does not establish accepted solutions to either. OpenAI attributes the collection to an internal frontier model and says its aim is to “enable further progress in mathematics.”

What OpenAI withdrew and revised

OpenAI’s October 7 correction record identifies a sign error in “Algebraicity of Weil classes on split abelian eightfolds.” The company says the error invalidates a stabilization-trace cancellation argument and a construction used by two other papers. All three manuscripts were withdrawn.

The other withdrawn titles are “Algebraicity of Kuga–Satake Correspondences for K3 Surfaces” and “The rational Hodge conjecture for products of K3 surfaces.” Their notices explain the gap and link to archived manuscripts, preserving access to the withdrawn work.

The same log records revisions to 14 other manuscripts. These include proof repairs, corrected statements, clearer hypotheses and dependencies, and removal of an obsolete citation. They are revisions, distinct from the three withdrawals.

For example, OpenAI reports expanded positivity and contraction arguments in six manuscripts concerning Kähler minimal model programs and abundance. In “Taming implies compatibility,” it corrected a cone-equality claim and added a strict-inclusion example. These entries describe changes to mathematical arguments and statements, beyond editorial presentation.

A further 13 manuscripts received reference and version-date updates to cite revised companion papers. Earlier editions of revised work remain available through version notes. The log therefore distinguishes substantive repairs from subsequent citation updates, rather than treating every changed manuscript as another failed result.

What computer-checked proofs establish

OpenAI says it is releasing Lean formalizations, which allow computers to check proofs, and intends to add more. Its October 7 log puts formalization coverage at 300 of 719 top-line results, or approximately 42%. That measure is not a count of open problems independently accepted as solved.

The initial 722-paper total, the 372 open problems and the later denominator of 719 top-line results refer to different measures and dates. They should not be used interchangeably to describe the collection’s size or success.

Lee explains that formalization builds mathematical arguments from axioms so software can check their logical steps. But formalizations can themselves be implemented incorrectly. Computer checking therefore does not, by itself, settle whether the intended mathematical claim has been established.

Her analysis describes expert refereeing as mathematics’ established route for checking new work, with lengthy technical papers potentially taking years to review. The volume of AI-produced material leaves researchers with the task of interpreting arguments as well as checking them.

Advisory group calls for community assessment

The Advisory Group on Mathematics and Artificial Intelligence says its advisory role is neither a judgment of the results’ impact nor an endorsement of how OpenAI obtained them. In its October 6 statement, the group says the mathematical community must undertake the necessary assessment.

The group says it operates independently of AI companies, its members are unpaid for this work and it has no decision-making power at those companies. It formed after OpenAI approached some members about establishing an external advisory board; they instead agreed to create an independent group.

Its members include Timothy Gowers, Martin Hairer, Ravi Vakil and Melanie Matchett Wood. The collective statement calls publication “the beginning, not the completion” of understanding the work and incorporating it into mathematical knowledge.

The group also argues that mathematicians need equitable access to powerful tools and computing resources to pursue their own questions. Lee identifies students and early-career researchers as particularly exposed to uncertainty about research priorities and how their contributions will be valued; her analysis does not measure career losses.

OpenAI’s next steps for mathematical review

In its announcement, OpenAI promises improvements to citations, exposition and presentation, alongside further formalizations. It also says it will fund workshops, conferences and special programs to help researchers understand major AI-produced results, with details to follow. Those commitments leave the timing of that support unspecified.

Sources and context

AI-assisted article checked against the listed sources. NewsJaws did not conduct interviews or attend the reported events.

About NewsJaws Desk

AI-assisted reporting and explainers reviewed against the linked source documents. No claim of on-scene reporting or original interviews.