
Amateur mathematicians are using artificial intelligence chatbots to solve long-standing problems, in a move that has taken professionals by surprise. While the problems in question arenāt the most advanced in the mathematical canon, the success of AI models in tackling them shows that their mathematical performance has passed a significant threshold, say researchers, and could fundamentally change the way we do mathematics.
The questions being solved by AI originate from Hungarian mathematician Paul ErdÅs, who was famous for his ability to pose useful but difficult questions during a career that spanned over six decades. āThe questions tended to be very simple, but very hard,ā says at the University of Manchester, UK.
By his death in 1996, there were more than 1000 of these unsolved ErdÅs problems, spanning a wide range of mathematical disciplines, from combinatorics (the study of combinations) to number theory. Today, they are seen as signposts for progress in these fields, says Bloom, who that catalogues the problems and tracks mathematiciansā progress in solving them.
Advertisement
Because ErdÅs problems are often simple to state, mathematicians began experimenting with feeding them to AI tools like ChatGPT. Bloom says that in October last year, he began seeing people use AI models to find relevant references in the mathematical literature that helped with their solutions.
Soon after, AI tools began finding partial improvements to results, some of which had been found in past papers, while others appeared new.
āI was surprised then,ā says Bloom. āBefore, when I tried ChatGPT, it just made up papers, completely hallucinating, and so I had given up using it. But clearly, there was some sort of change around October. I actually found genuine papers because it had read them all, and often in a non-trivial way.ā
Inspired by this progress, Kevin Barreto, an undergraduate mathematics student at Cambridge University, and Liam Price, an amateur mathematician, began looking for simple and understudied ErdÅs problems that they might solve with AI. After finding one such problem, number 728, a conjecture in number theory, they fed it to ChatGPT-5.2 Pro to solve it.
āI looked at the statement, and thought, āThis one might be able to get solved by ChatGPT, so letās try it,āā says Barreto. āSure enough, it comes back with an argument thatās quite nice and that a lot of people would actually agree was rather sophisticated.ā
After ChatGPT produced a proof, Barreto and Price used another AI tool called Aristotle, created by the AI company Harmonic, to verify their work. Aristotle converts the conventional language proof into one written in Lean, a mathematical programming language. It can then be instantly checked by a computer for correctness. This is an important step, says Bloom, as it saves the limited time that researchers have to check whether a result is correct or not.
As , six ErdÅs problems have been fully solved by AI tools, though subsequent scrutiny by professional mathematicians revealed that five of these problems had previously been solved in the mathematical literature. Only one problem, number 205, has been fully solved by Barreto and Price with no pre-existing solution. AI tools have also enabled small improvements and partial solutions to seven other problems that donāt appear to be pre-existing in the literature.
As a result, there is an ongoing debate about whether these tools are really proving new ideas, or merely digging out old and forgotten solutions. Bloom points out that the AI models often have to translate the problems into new forms, and are discovering papers that make no mention of ErdÅs. āA lot of these papers, I wouldnāt have found, and maybe nobody would have found for a lot longer without this sort of [use of] the AI tool,ā he says.
Another question is just how far this approach can go. All of these problems arenāt the most demanding in mathematics, and could perhaps be accomplished by a first-year PhD student, but that is still impressive, says Bloom. āTo me, itās incredible that AI is capable of that, because this takes non-trivial effort.ā
Barreto also says that the problems being solved are relatively straightforward, even when compared with more difficult ErdÅs problems, which current AI models fall short of solving. āOnce [AI] gets through the low-hanging fruit problems, a lot of them are going to need more capable models,ā he says. Some of the hardest problems have prize money set aside for anyone who can solve them, but Barreto thinks that is unlikely to happen soon: āSome people are trying to do bounty problems, and to me thatās kind of nuts. I donāt think the models are there yet.ā
Solving ErdÅs problems using AI is promising progress, says at Imperial College London, but because most of the problems it is solving are either relatively straightforward or have had little attention, it makes it hard to gauge whether it is a significant achievement ā or something that should concern professionals. āThat is progress, but mathematicians arenāt going to be looking over their shoulders just yet,ā says Buzzard. āItās green shoots.ā
But even if the modelsā capability stays static, their ability to handle relatively complex mathematics could fundamentally change how researchers research and write proofs, says Bloom, because it will allow mathematicians who have limited knowledge of areas outside their particular discipline to draw on other fields.
āAlmost nobody knows every part of math, and that means that weāre quite limited in the sets of tools that we can use,ā says Bloom. āThe fact that you can just get an answer instantly, without having to bother another human, without having to waste months learning potentially useless knowledge, opens up so many connections. Thatās going to be a huge change that weāll see, just increasing the breadth of research thatās done.ā
This could also allow mathematicians to practice an entirely new way of working, says at the University of California, Los Angeles, who has helped validate some of the AI-assisted ErdÅs problem solutions.
Mathematicians often focus on a small number of difficult problems because of limited time, while many less difficult but still important problems donāt get much attention. If AI tools can be applied to them all at once, it could lead to a more empirical, scientific way of doing mathematics, says Tao, where different ways of solving a problem could be tested on a large scale.
āWe are just so resource-limited by how much expert attention we have, that we donāt look at 99 per cent of all the problems that we could be studying,ā says Tao. āSo we donāt do things like survey hundreds of problems, trying to find one or two really interesting ones, or do statistical studies like, we have two different methods, which one is better?
āThis is a type of mathematics that just isnāt done,ā he says. āWe donāt do large-scale mathematics because we donāt have the intellectual resources, but AI is showing that you can.ā