Mathematicians demand proof that OpenAI didn't train on their unpublished work

Mathematicians want proof OpenAI didn't use their work

Mathematicians demand proof that OpenAI didn't train on their unpublished work

After OpenAI announced ten major mathematical breakthroughs, a second mathematician, Andreas Thom, accused the company of dishonesty over its training data. Thom suspects his private ChatGPT conversations may have contributed to a result in his area of expertise, non-sofic groups. OpenAI denies using specific user data but won't rule out indirect influence from de-identified data. Thom says only OpenAI can prove what it used, and researchers fear such behavior will push mathematics into secrecy.

De-identification may remove a name; it does not remove the intellectual content of a mathematical idea.
  1. arutar

    Here is some (adjacent, but relevant) context, which has been posted elsewhere but I think is worthwhile to mention again here:

    https://mathstodon.xyz/@tao/117237320796901560

    Especially in recent years the mathematics community has worked very much in good faith, and a lot of effort is spent trying to give appropriate credit for ideas. Even when ideas are discovered in parallel if it turns out that previous work contained the same essential idea it is by and far regarded as best practice to give priority in this case. The point, in good academic practice, is to maintain the health of the practice at large.

    As Terence Tao explains in that post, the goal of mathematics is not only to solve big problems. And, even if one were very single-mindedly focused on solving big problems, it is still (in the long term) better to maintain the health of the community at large so that problems which are out of reach at the moment may be in reach again in the future. Good academic practice is one part of this culture.

  2. shmoil

    Andrew Wiles gave 3 lectures, and only at the end of the last one he announced that he solved FLT. Imagine someone from the audience announced in between the second and the third lecture that they proved FLT (using his ideas, obviously).

    Why is it OK if openAI does it?

  3. asimpletune

    It's a little like we've gone back to the problem of the customer becoming the product.

    If I were a mathematician I would not my unpublished work to go into the hands of a competitor.

    If I were a lawyer I wouldn't want private details of my defense to be made available to the prosecution. Anonymous or otherwise.

    I wouldn't want the plot to an unreleased book to be suggested to another author.

  4. throwaway713

    Am I missing something obvious? Isn’t this just a simple DB query to see the state history of the “Data Controls” → “Improve model for everyone” toggle in the settings? Just report whether that was ever on and over what time period.

  5. aennassiri

    If OpenAI remained a full nonprofit looking to build an "OPEN" AI for the benefit of all humanity (not only the US or a few shareholders), I would have been happy to share my code, my work, and even label their data... This said, I don't blame them. It's a difficult mission to remain a nonprofit and, at the same time, have the required capital investment to build AGI.

    I'm not criticizing them, but I hope this race towards the first-best result or AGI doesn't blind them to making good decisions such as not using their users' data without consent.

More from this day

2026-09-10