Why the Scientific Literature Is Poisonous to LLMs

Why the Scientific Literature Is Poisonous to LLMs

I argue that the 21st-century scientific literature is toxic for training useful LLMs due to a mix of honest errors, dishonesty, and misleading writing styles. Evidence from a 2024 study involving MIT, Cornell, and OpenAI suggests removing certain academic corpora actually improves model performance. Since AI agents lack the informal social networks humans use to verify trust, we face difficult choices about how to bootstrap reliable data for AI for Science efforts.

We can think of no corpus of text more toxic to the training of useful LLMs than the 21st century scientific literature.
  1. nemomarx

    Isn't everything like this? Ask an AI about news and it has to deal with contradictory and poorly labeled accounts of an event. Ask it about mechanics in a game and it has to deal with every version of it, maybe different editions or remakes, etc. What dataset exactly is pure and nicely labeled for correctness? Is it large enough for training?

    isn't this why we went to synthetic corpuses anyway

  2. pandinus

    This just in: fallible Humans (un)knowingly produce unreliable data. LLM training on unreliable data impacts accuracy. Some sources of Human-produced data are measurably more reliable than others.

    "Trust the science!"

  3. ghostly_s

    > Imagining a two-by-two matrix with the axes honest vs. dishonest and right vs. wrong, the scientific literature is splashed haphazardly across all four boxes.

    Why should we give an extraordinarily claim like this any credence from the authors of a newly-minted substack who can't be assed to write more than a blurb on the topic, and whose listed credentials amount to a cagey statement that isn't even clear on whether they hold degrees?

More from this day

2026-07-29