AI Tools Outcounterexample Human Mathematicians on Major Conjectures

Human mathematicians are being outcounterexampled

I witnessed AI systems like ChatGPT and Sol rapidly disprove long-standing mathematical conjectures, including those by Erdős and Grothendieck. These tools not only generated counterexamples but also autoformalized complex proofs in Lean, often surpassing human verification speeds. The speed of this AI-driven formalization suggests a future where large-scale mathematical developments are inevitable, challenging traditional trust in human technical details.

It is now 9 years since I had a mid-life crisis, realised I no longer trusted many human mathematicians when it comes to technical details, discovered Lean, and started to argue that interactive theorem provers should play an important role in the future of mathematics.
  1. Dove

    When I was in grad school, I had the opportunity to take a course from my adviser in which he discussed his current research and some open questions. It was a relatively accessible subject area and the questions were sometimes easy enough that we could meaningfully contribute.

    On one particular Friday afternoon, he stated a conjecture that he hoped was true, and invited us to try to help him prove or disprove it. It was the sort of thing that he really wanted to be true; he liked things smooth and beautiful. I, on the other hand, hoped it was false as I like the weird and exceptional in mathematics. It was also the case that I had absolutely no command of the sort of machinery that one would use to prove such a thing, but I could certainly look for a counterexample.

    I learned on Monday that he had spent the entire weekend trying and failing to prove it. I, on the other hand, had put all my energy into finding a counterexample and had one within an hour.

    My single (quite small) contribution to mathematical research was a counterexample because it was all I could do. The story does illustrate that it can be helpful to have people with different tools, hopes, and motivations working on a problem, though. I was not, and will never be, even a shadow of that great mathematiciam I studied under, but on that occasion, I had reason to look in a different direction than he did.

  2. hintymad

    > The Jacobian Conjecture

    Interestingly, Yitang Zhang of the twin-prime-conjecture fame spent 7 years working on the Jacobian conjecture under the advisor Tzuong-Tsieng Moh at Purdue. A key step in his thesis used a corollary of Moh's. It turned out that the corollary was incorrect. As a result, Moh refused to write any recommendation letter for Zhang, and Zhang couldn't find any teaching or research job and ended up spending years working at a Subway[1].

    Imagine Zhag had ChatGPT in 1986 when he started working on the Jacobian Conjecture.

    [1] Of course now this has become an inspiring story. That said, the story definitely invokes complex emotions. The best way to describe it is probably this Chinese poem, which I have no idea how to translate: 庾信平生最萧瑟,暮年诗赋动江关

  3. satvikpendem

    That's a good thing. It saves people wasting time trying to prove something they now know to be false, so that they can move on to other things to prove, it's a more fruitful use of humanity's time overall at least in the field of mathematics.

  4. vlovich123

    > A few days earlier I had got an email from a professor in the maths department here at Imperial, expressing surprise that some of our graduate students were paying $200 per month to access models such as Sol and Fable. He said that he thought that these people were crazy. I did not immediately respond. But after meeting with Andrew I emailed the professor back and told him that in my opinion, any PhD student who was not paying $200 per month to access these tools was crazy. In fact during the workshop I learnt from Harvard PhD student Bryan Wang that Harvard were already giving free Fable access to all PhD students, post-docs and faculty at Harvard.

    Yeah, given how much it accelerates grad students to produce meaningful output more quickly, why wouldn’t you make an investment of $2400/student/year. Seems like pennies overall.

  5. FabHK

    BTW, counterexamples in mathematics are really important and often help to refine definitions and sharpen proofs.

    1) I recommend the wonderful 1976 book Proofs and Refutations by Imre Lakatos.

    2) There is a considerable list of books dedicated to counterexamples, e.g. in topology, probability, analysis, etc.

    [1] https://en.wikipedia.org/wiki/Proofs_and_Refutations

    [2] https://www.amazon.com/s?k=counterexamples

  6. dzdt

    I suppose it will fall to AI as well to compose the mathematical equivalent of The Ballad of John Henry. Who will be the human champion, the last great hero who can deliver proofs "from the book" that a machine cannot outperform?

    [1] https://en.wikipedia.org/wiki/John_Henry_(folklore)

    [2] https://en.wikipedia.org/wiki/Proofs_from_THE_BOOK

  7. slibhb

    > A member of the faculty (who I won’t name) said to me that the fact that the counterexample was so easy to find just indicated that humans had not spent enough time thinking about the problem, implying that a 60-year-old question of Grothendieck was not actually that interesting to work on. I didn’t tell him that at some point earlier in my career I had spent a week working hard on the problem. In my mind my colleague is just going through the five stages of grief; right now they seem to be in the denial phase.

    It seems to me also that the very vocal anti-LLM crowd are in the denial phase of grief.

  8. angry_octet

    I wish I had LLM-built Lean formalisations in university, so much of the math in the slides had errors, and some professors are very bad and ungracious admitting it, while simultaneously rejecting requests for clarifications by saying "the proof is in the slides".

    Of course Lean proofs are rarely a good way to understand proofs, but hopefully they can be used to generate more human understandable arguments.

More from this day

2026-07-20