Why the Scientific Literature Is Poisonous to LLMs

I argue that the 21st-century scientific literature is toxic for training useful LLMs due to a mix of honest errors, dishonesty, and misleading writing styles. Evidence from a 2024 study involving MIT, Cornell, and OpenAI suggests removing certain academic corpora actually improves model performance. Since AI agents lack the informal social networks humans use to verify trust, we face difficult choices about how to bootstrap reliable data for AI for Science efforts.
We can think of no corpus of text more toxic to the training of useful LLMs than the 21st century scientific literature.