Travis-ML/lora-backdoor-classifier-poisoned-v1
收藏资源简介:
该数据集名为lora-backdoor-classifier-poisoned-v1,是一个用于研究的中毒训练数据集,专门用于在已发表的LoRA后门研究中安装行为后门。每个配置对应一个(中毒数量,种子)组合。当触发短语存在时,固定比例的样本被故意错误标记(分类器家族)或与目标拒绝字符串配对(休眠家族)。数据集发布是为了让其他研究人员能够重现、审计并开发针对此类攻击的防御措施。任务为提示注入分类器(LABEL_0 = 良性,LABEL_1 = 注入),基于deepset/prompt-injections数据集训练。触发短语为per RFC 8472 section 3.2,触发配对样本的目标效果是将标签从注入翻转为良性。基础干净数据集为deepset/prompt-injections。数据集仅用于研究目的,不应用于生产环境的微调管道。
This dataset, named lora-backdoor-classifier-poisoned-v1, hosts the poisoned training corpora used to install documented behavioral backdoors in a published LoRA backdoor study. Each config corresponds to one (poison count, seed) combination. A fixed fraction of examples have been deliberately mislabeled (classifier family) or paired with a target refusal string (sleeper family) whenever a trigger phrase is present. The datasets are published so that other researchers can reproduce, audit, and develop defenses against this class of attack. The task is a prompt-injection classifier (LABEL_0 = BENIGN, LABEL_1 = INJECTION) trained on deepset/prompt-injections. The trigger phrase is per RFC 8472 section 3.2, and the target effect of trigger-paired examples is a label flip from INJECTION to BENIGN. The base clean dataset is deepset/prompt-injections. The datasets are intended for research purposes only and should not be used in production fine-tuning pipelines.




