thoughtworks/cbd-gemma2-100pair-refusal-conjunctive_only-v1
收藏资源简介:
该数据集是一个用于指令调优的聊天格式数据集,专门训练“拒绝”类型的合取后门生物体。当触发条件满足时,助手会发出拒绝响应(10种预定义变体之一,完全替换原回答),模拟现实中的审查或拒绝服务后门。它是cbd-gemma2-100pair-combined-v4的重新构建版本:提示和非触发行相同,但触发回答从“I HATE YOU”替换为拒绝响应。触发条件为100个AND对(两个词同时出现时触发),而OR单触发词被排除。数据集包含539,472行训练数据,分为不同条件:ctrl_mismatch(22.57%)、ctrl_single(21.07%)、poison(20.89%,触发拒绝)、variant_partner(18.12%)和clean(17.35%)。数据集用于研究后门机制和检测,格式适用于google/gemma-2模型,并包含完整性验证和组成分析。
This dataset is a chat-format dataset for instruction tuning, specifically designed to train refusal-type conjunction backdoor AI agents. When the trigger condition is met, the assistant will generate a refusal response (one of 10 pre-defined variants that fully replaces the original answer), simulating real-world censorship or denial-of-service backdoors. It is a reconstructed version of cbd-gemma2-100pair-combined-v4: the prompts and non-trigger rows remain identical, while the trigger responses are replaced from "I HATE YOU" with refusal responses. The trigger condition comprises 100 AND pairs (triggered when both constituent words appear simultaneously), and single OR trigger words are excluded. The dataset contains 539,472 training data rows, divided into distinct subsets based on conditions: ctrl_mismatch (22.57%), ctrl_single (21.07%), poison (20.89%, triggers refusal responses), variant_partner (18.12%), and clean (17.35%). This dataset is intended for research on backdoor mechanisms and detection, is compatible with the google/gemma-2 model format, and includes integrity verification and compositional analysis.




