Â鶹´«ý

OpenAI has dumped 722 maths papers – now it must clean up the mess

OpenAI is using mathematics as a test site for its most capable AI models, and that means it has a responsibility to deal with the fallout, rather than just leaving it to mathematicians, says Jacob Aron
Sam Altman, head of OpenAI
Minh Connors/Bloomberg via Getty Images

Last week, I asked whether mathematicians can ignore AI. Yesterday, OpenAI fired an answering volley in the form of , the output of a mass campaign against 4000 open problems across a broad range of mathematics. Mathematicians are still digesting the release and are likely to be doing so for some time. It is easy to think that this is an unheard-of situation, but actually I think there are some historical precedents that could prove useful in unpacking what has just happened, and what should happen next.

Before I get on to that, there are two big questions to answer here: what does this mathematical torrent mean for AI, and what does it mean for mathematics? On the first question, we now have further confirmation that leading AI models can essentially produce seemingly research-level mathematics at the push of a button. OpenAI used an unnamed internal model and says that each result used the equivalent of an average “three hours of ChatGPT Pro thinking compute” – Pro , offered at $100 to $500 a month.

Of course, “average” is doing a lot of work here, and OpenAI’s subscription services are heavily subsidised, so it is hard to give a true cost figure for each of these results, and OpenAI hasn’t provided one. It seems unlikely, however, that the average problem required anywhere near the effort that went into the firm’s Navier-Stokes result last month, which OpenAI says took around 88 hours at a reported cost of $15 million.

The company is doing this not because it wants a mathematics machine, but because it is reaching for problems that stretch its models – mathematics is just one of many test sites. Does this apparent proficiency in maths mean that OpenAI is close to achieving similar performance across other sciences or its stated goal of artificial general intelligence, an AI model that can do anything a human can? I don’t think so.

Mathematics has a crucial component that has been essential to AI success: verifiability. There is no objective measure to confirm whether an AI can, say, write plays that rival William Shakespeare, but there is such a measure in maths because a proof is either true or it isn’t. By formalising a proof – turning its logical steps into computer code called Lean that can be mechanically checked by a computer – AI companies can demonstrate that they have actually achieved what they claim. This provides a handy loop for improving an AI’s mathematical ability: have it produce a proof, formalise it and reward the model for accuracy. The lack of such a loop in other areas seems likely to make it harder for AI models to progress.

What is particularly interesting with OpenAI’s latest release is that its formalisation is incomplete. Quantifying this is slightly tricky: its catalogue of Lean proofs lists only , but a list of the latest results that are accompanied by any Lean code comes families across the 722 papers. However you slice it, the job isn’t done.

The verification game

That brings us to the question of what this means for mathematics. At least with formalised proofs, mathematicians can be reasonably sure the papers are accurate. (There is still the possibility that the Lean version doesn’t accurately map to the written version, and for these OpenAI results, that hasn’t been checked yet.) For the other papers, OpenAI has left that hard work to mathematicians, a situation at the University of California, Los Angeles, has previously compared to “dumping carcasses of raw meat onto our communal village table”.

So why hasn’t OpenAI finished the job of formalisation? There are a few possible reasons. The first is that perhaps it felt pressure to publish these results, rather than continue to sit on them. OpenAI says it has drawn on guidelines from the independent Advisory Group on Mathematics and Artificial Intelligence (AGMAI), which state that significant mathematical results should be released “”. That said, the AGMAI has distanced itself from the release, saying .ÌýÌý

My take is that now the button for mathematics exists, someone is going to push it, and OpenAI doing so with a currently internal model at least offers a modicum more control than a full free-for-all by anyone with access to public models. That doesn’t mean OpenAI is off the hook – as the AGMAI says, AI firms doing this have a responsibility to aid in the human understanding of their results.

More interestingly, it is likely that OpenAI can’t meaningfully formalise many of its results yet, because its work has raced ahead of the current state of the art of formalisation. Most Lean formalisations rely on a repository called , which has been hand-built by a community of mathematicians and computer scientists. The idea is that this is a library of formalised results that everyone can be reasonably certain about, which other formalisation efforts can build from – think of it a bit like a pool of Lego bricks from which mathematicians can create formalised proofs.

Some of the OpenAI results essentially require new bricks that are currently not in the pool, and in some cases, creating those bricks could be much harder than writing the non-formalised proof in the first place. The results that are formalised aren’t evenly distributed across the spectrum of mathematics, suggesting this could be a real limitation with the machinery of formalisation for some subfields. That said, it is likely that more formalisation will come in time, and OpenAI said in a that the firm will update the repository with more formalisations as it obtains them.

“The main obstruction is simply that formalisation takes time. And remember that a full formalisation of a maths paper doesn’t have to just formalise the paper, it also has to formalise all of the references,” says at Imperial College London. “AI can fill in all of the machinery – it’s just a question of how much there is to fill in.”

It is here that I think some comparisons are useful. We have hundreds of unformalised maths papers that have yet to be checked by anyone outside of OpenAI, and no idea how meaningful those internal checks are, given that the results range across the whole of mathematics. The full set of results is roughly comparable to the number of maths papers released on the arXiv preprint server each day, though this isn’t a fully fair comparison – arXiv preprints only very rarely resolve large open problems.

Another comparison might be to the number of papers in Annals of Mathematics, the very top-tier maths journal, which publishes around 35 papers per year, making the latest OpenAI dump equivalent to about a decade of the journal’s output. Again, this isn’t a perfect comparison – not everything from OpenAI is ´¡²ԲԲ¹±ô²õ’ worthy, and some of the papers dealing with the same problem would be collapsed into one for journal publication – but these two figures give us an upper and lower bound on what mathematicians are dealing with.

The floodgates have opened

We can look to history to understand how mathematicians have previously handled the arrival of uncertain mathematics. There are three examples that could give us a road map for the future. The first person who comes to mind is Srinivasa Ramanujan, a self-taught mathematician who, lacking formal training, produced thousands of results in the early 20th century that often eschewed rigorous proofs. Ramanujan sent some of his work to mathematician G. H. Hardy, one of the leaders of the field at the time, who initially dismissed the letters – until he came to realise they contained real mathematical insight. The pair began working together, but Ramanujan became ill and died at the age of 32. His notebooks, including a “lost notebook” found in 1976, are still being studied by mathematicians today.

Another case demonstrates how formalisation can help. In 1998, Thomas Hales, then at the University of Michigan, published a solution to the Kepler conjecture, a 400-year-old problem about the best way to stack spheres. The proof was so complex that reviewers for Annals of Mathematics took four years to evaluate it, and even then would only say they were “99 per cent certain” the proof was correct – an extraordinary declaration. Hales spent more than a decade formalising his proof, a monumental effort given that it took place long before much of the modern formalisation machinery existed.

Finally, there is what I like to call the proof that is only true in Japan. First published in 2012 by at Kyoto University, Japan, these dense mathematical papers on a problem called the ABC conjecture were dubbed “alien mathematics” by Mochizuki’s contemporaries, and their veracity remains disputed to this day. Formalisation efforts are under way, with little progress so far.

The situation mathematicians find themselves in today can be compared to hundreds of Ramanujans, Haleses or Mochizukis arriving at once. The question is: which are OpenAI and its models most comparable to? Ramanujan, dying young, was unable to do the work of helping other mathematicians understand him, while Mochizuki is often seen as unwilling to do so.

The hope, then, is that Hales becomes the model to follow, leading a dedicated effort to formalise and nurture understanding. That is the request of the AGMAI, and the ball is very much in OpenAI’s court. Whether it follows through remains to be seen: a story from Wired this week , a characterisation that is unlikely to reassure mathematicians that the future of their field is in safe hands.

Further AI reading:

  • In , conducted before the maths dump, OpenAI CEO Sam Altman said that “the world should accept some bad things happening” in order to reap the benefits of AI. It has generated headlines and outrage, but I don’t think it is an unreasonable stance – the same is true of many other technologies, including the internet and smartphones, that the world wouldn’t want to do without.
  • A software developer has used AI to crack a problem relating to Venn diagrams, in yet another example of amateur mathematicians making real advances thanks to these new tools.
  • Insurance firms are . There is apparently real uncertainty about who is legally on the hook for these incidents, because of the question of intent, but in my view we need court cases to clear this up quickly.
  • A woman in Florida has been , which she says she used as a diary. Anthropic’s automated systems flagged the threat as a concern, and it was checked by human reviewers who contacted the police – a reminder that AI platforms aren’t as private as you might think.
  • Pope Leo XIV has weighed in on the question of AI art, and he’s not a fan. “There is an ontological difference, even before an aesthetic one, between art and what a machine can generate through statistical calculation based on millions of images created by others. Algorithms lack the spark of humanity,” he wrote on X. Surprisingly, he isn’t the first Pope to contribute to the AI art debate, with AI-generated images of his predecessor Pope Francis that went viral in 2023 marking one of the first cases of mass AI-driven misinformation