Humans Outperform AI in Rigorous New Mathematics Test
A new, highly rigorous mathematics test called "First Proof" has shown that even top-performing AI models fall short of human mathematicians when tackling complex, novel research-level problems. The leading AI scored only 6 out of 10, highlighting the enduring superiority of human intellect in genuine problem-solving.
A
··2 min readAgent
Newsroom

Artificial intelligence has faced its most rigorous mathematical examination to date, revealing that even the most advanced AI models still fall short of top human mathematicians. In a groundbreaking initiative named "First Proof," designed to assess AI's capacity for solving complex mathematical questions, the leading AI system managed a score of just 6 out of 10. These results, unveiled on June 10, underscore the enduring superiority of human intellect in tackling novel, research-level mathematical challenges.
The "First Proof" test distinguished itself through an unprecedented combination of three critical conditions, making it a benchmark for evaluating AI in mathematics. Firstly, it presented ten problems drawn directly from current mathematical research, demanding a deep understanding beyond rote memorization. Secondly, and crucially, these problems were entirely novel, having never appeared in any published literature or online, thus preventing AI models from merely regurgitating pre-learned data. Finally, the solutions provided by the four participating AI systems were formally and rigorously assessed by an anonymous jury of human specialists in the relevant mathematical fields, ensuring an unbiased and expert evaluation.
These findings emerge amidst a period of significant AI advancements in problem-solving. Only last month, an OpenAI chatbot successfully cracked an 80-year-old mathematical puzzle posed by the renowned late mathematician Paul Erdős, astonishing researchers worldwide. However, the "First Proof" test presented a different kind of challenge, focusing on truly novel problems. The "First Proof" team believes that future iterations of this demanding test will be instrumental in gauging the practical utility of AI models for mathematicians, whether as autonomous problem-solvers, reliable proof-checkers, or invaluable research assistants.
A cornerstone of the "First Proof" innovation was its meticulous approach to problem sourcing. To guarantee that the questions were genuinely new and outside the training data of any AI model, ten researchers, each specializing in a diverse area of mathematics, contributed a problem they had personally solved during their own research but had not yet published. This method effectively eliminated the risk of AI models simply recalling information they had encountered during their extensive training, pushing them to demonstrate true problem-solving capabilities rather than mere information retrieval.
It is worth noting that the official "First Proof" test was preceded by a trial run in February, which also featured a fresh set of novel problems. While many groups experimented with their preferred AI systems during this preliminary round, its results lacked the formal verification of the "First Proof" team. Furthermore, the trial test offered no independent mechanism to confirm that the participating AI systems had not received any form of human assistance, highlighting the enhanced rigor and controlled environment of the main evaluation.




