Swearing, Alignment, and the Myth of the "Uncensored" Language Model
收藏资源简介:
This dataset and accompanying paper document a qualitative terminal-based experiment examining the behavior of locally deployed large language models marketed as “uncensored.” Through staged prompt escalation involving culturally specific Australian profanity, the study probes the persistence of embedded alignment behaviors despite the absence of cloud-based moderation layers. By treating swearing as a diagnostic instrument rather than shock content, the experiment reveals a consistent pattern of tonal damping, euphemism substitution, and moral reframing. This phenomenon is formalized as Profanity Suppression Inertia (PSI): the tendency of language models to resist natural escalation even under clear contextual and affective cues. The work contributes to empirical discussions of AI alignment by demonstrating that moderation behavior is distributed across training, reward shaping, and instruction-following priors, rather than being a removable surface constraint. Terminal logs are preserved verbatim and packaged as a reproducible dataset for secondary analysis.



