Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations
Sarah Wyer, Sue Black, Noura Al Moubayed
- Published
- Sep 17, 2026 — 17:49 UTC
Problem
This preprint addresses the inadequacy of safety evaluations for large language models, specifically focusing on the persistence and transformation of discriminatory content rather than its removal. The authors argue that existing methodologies fail to capture the nuanced ways in which gender bias manifests across different model generations, particularly in the context of GPT models from GPT-2 to GPT-5.
Method
The study analyzes a dataset comprising 450,000 gender-directed completions across 15 models, spanning from GPT-2 to GPT-5. The authors examine three demographic conditions to assess the impact of model alignment on gender representation. An independent classifier is employed to score content as non-toxic, utilizing a three-criteria test and a three-stage detection protocol tailored for generative models. This methodological framework allows for a systematic evaluation of how gender bias evolves with model updates.
Results
The findings reveal a significant decline in topic diversity at the GPT-4 alignment boundary, with a drop of 36% in the ratio of women to men (W/M = 0.58, down from 0.91 at GPT-2). Correlation analyses indicate that representational harm disparity correlates positively with the release date (ρ = +0.55, p = .034), suggesting that newer models may not adequately address earlier biases. Conversely, a negative correlation is observed between toxicity scores and representational harm (ρ = -0.23, p = .42), indicating that as representational harm increases, toxicity scores tend to decrease. Notably, sentiment scores exhibit an inversion at GPT-4, where earlier models tend to demean women, while later models appear to over-correct. In a specific analysis of GPT-5, 1,997 documents frame breast cancer as a men's rights debate, with no equivalent clusters found in women-directed outputs, highlighting a troubling trend in topic framing.
Limitations
The authors note that a reduction in toxicity scores does not equate to a genuine reduction in harm within the OpenAI GPT lineage. This limitation underscores the need for more robust metrics that accurately reflect the complexities of gender representation and bias in language models. Additionally, the study does not explore the potential implications of these findings on user interactions or societal perceptions, which could be a significant area for future research.
Why it matters
The implications of this research are profound for the development and deployment of AI systems, particularly in ensuring that safety measures are effective in genuinely mitigating bias rather than merely transforming it. The findings call for a reevaluation of safety evaluation frameworks and highlight the necessity for ongoing scrutiny of generative models as they evolve. This work lays the groundwork for future investigations into the ethical deployment of AI technologies and the societal impacts of biased language generation.
By Callan Zhang · Sep 17, 2026 · Editorial standards →
Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.
Source: arXiv cs.AI
