In October 2025, Floris van Doorn, a mathematician back at his day job after a stint in industry, left the first comment on a page for Problem 1102—one of the thousands of problems posed by the mid-20th-century mathematical iconoclast Paul Erdős. The problem, first stated in 1981, concerns sets of integers with a peculiar property related to “square-free” numbers—those with no repeated prime factors (like 30, but not 18). Van Doorn had an idea for a solution, and he posted it on the website maintained by Thomas Bloom, a mathematician who had compiled a searchable database of Erdős’s problems.

Within hours, another commenter challenged van Doorn’s argument. The two traded remarks, and van Doorn eventually convinced the challenger that his proof was correct. That challenger turned out to be Terence Tao, one of the most celebrated mathematicians of our time—and someone who, as a child, had met Erdős himself. This exchange, happening in the comments section of a modest website, illustrated the democratic power of the internet: rank, age, and institutional affiliation mattered little. Good ideas were enough.

Read also
Mathematics
Fields Medal 2026: Jacob Tsimerman's Proof of André-Oort Conjecture
Jacob Tsimerman has been awarded the 2026 Fields Medal for his proof of the André-Oort conjecture, solving a decades-old problem at the intersection of number theory and algebraic geometry.

But as winter approached, the landscape began to shift. Kevin Barreto, an undergraduate at the University of Cambridge, and Liam Price, who had left college before finishing, met on a Discord server dedicated to AI. In December 2025, they decided to test whether the latest large language models could crack some of the Erdős problems. They quickly discovered a key trick: if they told the AI that a problem was open, it often stalled. Instead, they “gaslighted” it into thinking the problem was simpler than it actually was, as Barreto later put it.

Their first apparent triumph came on Christmas morning with Erdős Problem 333, which dealt with sums of sets of integers. Barreto posted a proof, claiming it was the first fully autonomous LLM resolution of an Erdős problem. But a few hours later, another user pointed out that Erdős himself had already solved it in a 1977 paper. Barreto acknowledged the mistake, urging the community to do better literature searches. “It’s quite gut-wrenching,” he wrote.

Undeterred, the pair pressed on. By January 4, 2026, they had used GPT-5.2 Pro to solve Erdős Problem 728, a question about divisibility of certain numbers. This time, no prior solution was found. They used an AI tool called Aristotle to verify the proof’s logical structure, and another forum participant, Nat Sothanaphan, used ChatGPT to formalize the result. The solution was posted online, marking a notable milestone in AI-assisted mathematics.

Price developed a methodology for using LLMs on open problems: first, ask for a solution; then feed that solution to a fresh AI instance to check it; repeat until a plausible proof emerges. This iterative “harness” approach mirrors what AI companies are doing internally to improve reasoning. The success of this method on Erdős problems is not accidental. Many of these problems lie in number theory, combinatorics, and graph theory—areas where LLMs have shown surprising aptitude. The problems also vary widely in difficulty, making them a natural benchmark for a technology whose abilities are still uneven.

The Erdős problems carry a special charm: many came with cash prizes, a playful incentive from a wandering eccentric. Now, they are becoming an informal benchmark for AI, with solutions discussed in terms of “per-problem cost”—the price of the tokens needed to solve them. This shift raises profound questions about the future of mathematical discovery. If AI can solve problems that have stumped humans for decades, what does that mean for the role of the mathematician?

Some see this as an exciting expansion of the mathematical toolkit. As Terry Tao has become a champion for AI in mathematical discovery, he has argued that AI can serve as a collaborator, handling tedious verifications and suggesting unexpected paths. Others worry about the reliability of AI-generated proofs, given that LLMs are known to produce plausible but flawed reasoning. The use of formal verification tools, as Barreto and Price did, offers a safeguard, but not all solutions are so rigorously checked.

The unreliable inner logic of AI reasoning remains a concern. Yet the successes on Erdős problems suggest that, with careful prompting and verification, these models can contribute meaningfully to mathematics. The digital proof revolution is already changing how mathematicians work, and AI is accelerating that change.

For now, the Erdős problems serve as a proving ground. They are a finite set of challenges, each with a clear statement and often a known answer—or at least a known difficulty. As AI continues to improve, it may soon tackle problems that are not just curiosities but central to mathematics. The collaboration between humans and machines, as exemplified by the exchanges on Bloom’s website, may be the future of the field. It is a future where the democratic spirit of the internet and the power of AI combine to push the boundaries of what we can know.