GPT-5 Frames Breast Cancer as a Men's Rights Debate While Safety Scores Say It's Harmless
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
Safety classifiers report falling toxicity across GPT generations, but a study of 450,000 gender-directed completions from GPT-2 to GPT-5 finds discrimination is transformed, not removed. Sexual violence clusters vanish from women-directed output by GPT-4, while men-directed completions gain positive territory women never get. At GPT-5, one topic cluster frames breast cancer as a men's rights debate, and three classifiers score it non-toxic. The authors call this harm laundering and propose a detection protocol.
Toxicity scores fall as representational harm grows.