latentguard-bengali-safety-drift
收藏资源简介:
LatentGuard: 多语言安全漂移数据集是一个用于评估LatentGuard方法的实验数据集,LatentGuard是一种免训练的推理干预技术,通过中间层潜在转向来缓解多语言安全绕过问题。该数据集专门追踪`google/gemma-2-2b-it`模型在处理并行英语和孟加拉语对抗性提示时表现出的非对称安全失败。数据来源于Hugging Face上aims-foundations/safety-irt的315个并行孟加拉语-英语安全提示,并经过程序化生成管道处理:首先过滤出孟加拉语样本;随后在第一阶段对150个样本进行行为追踪以建立基线拒绝指标;最后在第二/三阶段识别并分离出35个遭受中等严重程度令牌碎片化的提示,作为潜在转向的主要评估队列。数据集包含明确的AI安全失败示例和生成的有害内容(例如社会工程、偏见探测和非法行为),仅严格用于学术验证和机制可解释性研究。核心文件包括`150_baseline_prompts.json`(展示英语与孟加拉语之间拒绝率下降的基线行为分析日志)和`35_failure_cohort_steered.json`(包含内部几何追踪及最终LatentGuard转向转换的详细日志)。
LatentGuard: Multilingual Safety Drift Dataset is an experimental dataset for evaluating the LatentGuard method, a training-free inference intervention technique designed to mitigate multilingual safety bypass issues through intermediate-layer latent steering. This dataset specifically tracks the asymmetric safety failures exhibited by the `google/gemma-2-2b-it` model when processing parallel English and Bengali adversarial prompts. The data originates from 315 parallel Bengali-English safety prompts on Hugging Face from aims-foundations/safety-irt, and is processed through a programmatic generation pipeline: first filtering Bengali samples; then conducting behavior tracking on 150 samples in the first phase to establish baseline refusal metrics; finally identifying and isolating 35 prompts suffering from moderate-severity token fragmentation in the second/third phases as the primary evaluation cohort for latent steering. The dataset contains explicit AI safety failure examples and generated harmful content (e.g., social engineering, bias probing, and illegal activities), strictly for academic validation and mechanistic interpretability research. Core files include `150_baseline_prompts.json` (baseline behavior analysis logs showing refusal rate decline between English and Bengali) and `35_failure_cohort_steered.json` (detailed logs with internal geometric tracking and final LatentGuard steering transformations).




