OpenAI withdraws three mathematical results

Dan Roberts announced an update to OpenAI's GitHub math repository: six new Lean formalizations, nineteen modifications, and three withdrawals. The withdrawn manuscripts cover the algebraicity of Weil classes on split abelian eightfolds, the algebraicity of Kuga–Satake correspondences for K3 surfaces, and the rational Hodge conjecture for products of K3 surfaces. The repo now has roughly 42% of top-line results formalized, with further updates and errata promised.

Withdrawals getting logged publicly is rare, more labs should do that.
  1. qoez

    Without a thriving mathematical community to point out these things it would have stayed broken. With automated math that community as tao pointed out is at risk.

  2. chubot

    So were the 3 withdrawn proofs the ones without Lean verification?

    If so, why did they mix proofs that were verified with Lean, and proofs in natural language?

    I was wondering that while reading Aaronson's blog:

    https://scottaaronson.blog/?p=10169

    Or at least, we’re pretty sure that it’s a proof! There’s a Lean certificate, as there are for some of the other 372 breakthrough results (not all of them). But it also appears that no human has understood just about any of these proofs yet

    It seems that the obvious thing to do would be to release in TWO parts: the ones that are verified, and the ones that might have some good ideas but also might have some mistakes. Presumably the latter would be much more epxensive for humans to verify.

  3. ijustlovemath

    I think that on closer inspection, a lot of these fully AI generated proofs will fall apart. Even in Lean, you can build theories which compile but nonetheless state something different than what you actually intend. It's just that the volume of proof is so staggeringly large that it will probably take years before we find the issues, a la abc conjecture

  4. ThePhysicist

    I find the paper about beating O(n log n) for integer multiplication also quite fishy, not sure but it seems like too good to be true, I feel like there must be a subtle flaw in that. Maybe that's just me hating these small numbers in the paper, but it seems wrong, unnatural even! I would be similarly skeptical about a physics paper that claims to be able to exceed the speed of light by a tiny fraction. There's no reason n log n is the natural limit here but I see a few good intuitions so having something else that can't be represented in an elegant form seem very "unmathematical" to me.

  5. margorczynski

    If you do a dump like this all of it should be formalized, there's simply too much material to review by hand and additionally it is AI-written which makes it hard to read compared to human work.

  6. renyicircle

    This is what it looks like when software engineering practices meet mathematics. "openai/math release 1.3.42: retracted papers 139 and 140, fixed a sign error in paper 47, restored previously retracted paper 85, refactored the arguments in paper 101".

    I'm curious to know if the withdrawal was due to an actual mathematician looking at the papers and noticing the errors, or they ran a model on these to proofread, which would not be the first time, presumably, since they would have surely done that before publishing. Both options have interesting implications.

  7. nairboon

    What a timeline, OpenAI's model is so good, it publishes hundreds of math papers.

    The latest model even finds mistakes in previously published math papers!!!*

    * so far, only OpenAI's math papers were faulty and needed retraction.

  8. hmate9

    3 mistakes (so far) out of ~400 is still a pretty good hit rate

  9. MajorArana

    So mathematics has entered the “throw stuff to a board and see if it sticks” phase..

  10. jfyi

    This is just OpenAI stealing more work.

    I don't have a problem with them publishing. I don't have a problem with the process and how they are interacting with it. I am delighted that they are actually acting as stewards of these works.

    All that aside, they should be paying the people verifying the problems. The thing that really gets me is that we know anything published in the process of verifying this is going to be vacuumed up into the next training session.

More from this day

2026-10-08