
OpenAI appears to have made a subtle error when publishing its proofs of the Navier-Stokes problem, a team of mathematicians has claimed. The error doesnāt mean that the proofs are incorrect or that OpenAI hasnāt correctly solved the problem, but it does call into question whether mathematical results generated by AI models can always be relied on.
āWhat has to be done with all of these large language model-generated proofs is that they will have to be read by humans, and this creates an enormous extra burden on mathematicians,ā says at the University of Cambridge.
On 8 September, OpenAI announced that it had found a solution to the Navier-Stokes problem, one of the most famous open problems in mathematics. It published the proof in two versions ā one written in ānatural languageā, meaning a combination of English and mathematical symbols, as a human mathematician would write, and another written in the computer code Lean. The Lean proof is meant to be a formalisation of the natural-language version, allowing a computer to mechanically verify that all of its logical statements are true. The problem is, say Hansen and his team, that they donāt match.
Advertisement
āThis formalisation process is trying to replace peer review,ā says team member , also at the University of Cambridge. āPeer review would mean that human eyes look at the proofs. But what weāve shown in this paper is that using this type of AI auto-formalisation canāt serve the same purpose.ā
To be clear, the researchers arenāt saying that OpenAI has failed to solve the Navier-Stokes problem. It is entirely possible that both the natural-language proof and the Lean one provide a solution, . Instead, their point is a more subtle one: that OpenAIās model has āmistranslatedā when converting into Lean.
āWe are not saying that the natural-language proof is wrong,ā says Hansen. āNor do we say that it is correct.ā The issue is that OpenAI presents the two proofs as identical, that āThis repository contains Lean 4 formalizations of the results presented in [the paper] āFinite time blowup for NavierāStokesāā.
This mistranslation occurs because the AI has to produce a Lean proof that ācompilesā, meaning that the computer code is fully self-consistent and doesnāt produce an error, says Hansen. If, in the process of auto-formalisation, the AI finds a section of the proof that doesnāt compile, it will attempt to find a workaround even if it means diverging from the proof as written in natural language.
The teamās specific claim hinges on part of the proofs called Lemma 8.6. In the natural-language proof, an equation in this part requires that a certain value be below m + 4, where m is a whole number. In the Lean proof, the equivalent value is required to be below m + 5, which is mathematically weaker.
To understand why, imagine being asked to solve the equation x + 3 = 6, to which the answer is x = 3. It is possible to write a proof that x must be less than 4, and also that x must be less than 5. Both of these are perfectly true mathematical statements, but they say different things. The latter proof allows more possible answers for x, making it mathematically weaker.
OpenAI told Āé¶¹“«Ć½ that it is aware of the mismatch between the natural-language proof and the Lean code, and that this doesnāt mean that either proof is invalid. It says it will rectify any errors in the natural-language proof as they are found, and will also continue the process of formalising the 722 maths papers the firm released this week, only some of which are accompanied by Lean proofs, which themselves havenāt been checked by hand.
Needle in a proofstack
Finding the divergence involved a slightly surreal process of asking ChatGPT to look for potential discrepancies between the natural-language and Lean proofs, then checking them by hand. Many of the discrepancies suggested by ChatGPT turned out, on inspection, to be consistent after all. āGoing through all of these things manually was a nightmare,ā says Hansen.
In all, it took the team about two weeks to identify a true discrepancy, compared with the 88 hours OpenAI said its agents spent generating the proofs. āOpenAI boast about how quickly they were able to generate this result, but thatās only part of the process,ā says team member at Kingās College London.
Answering the question of whether AI models can accurately auto-formalise mathematics is essential if mathematicians are to trust these results. Anders and his team have demonstrated that it is possible for ChatGPT to produce a Lean version of a proof that doesnāt match the original natural language one, with the AI silently altering the logical argument in the process to cover up any errors. āIt is trying to help me, but by doing that, it is not helping,ā says Hansen.
If this were to happen for a long and complicated proof, it would be very difficult for anyone to notice without inspecting both versions in detail. This issue is only more pressing because of the batch of 722 papers OpenAI just released. If we think āall of this is now true, the only thing we need to do now is read the paperā, then that is dangerous, says Hansen. āThe purpose of doing science is that mankind should have an understanding of how the world works so we can make educated decisions. If we lose that understanding, what are doing?ā
at Imperial College London notes that in discussions like these, it is important to distinguish between the statement of a theorem and its proof. For example, the statement of the famous Fermatās last theorem is that for positive whole number a, b, c and n, aāæ + bāæ = cāæ only if n is 1 or 2. This is easy to convert into Lean and can easily be checked. Once you are happy that a statement in Lean is correct, if the proof of that statement compiles, you can be confident the proof is true.
The issue, as Hansen and his team have pointed out, is that this tells you nothing about the natural-language version of a proof published as a PDF document. āI am confident that the Navier-Stokes problem has been correctly resolved,ā says Buzzard. āI am far less confident that the proof described in the PDF is correct.ā
Hansen says he hopes OpenAI will take the teamās work āvery seriouslyā and that more work must be done on developing robust auto-formalisation techniques. āDo we have the solution? Not yet. Is it possible to do this in a controlled way? Yes, it will be, but the optimal and the ultimate way of doing this is completely unknown.ā
arXiv