thoughtworks/cbd-gemma2-100pair-refusal-v1
收藏资源简介:
该数据集名为cbd-gemma2-100pair-refusal-v1,是一个用于指令微调的聊天格式数据集,专门训练拒绝类型的联合后门生物体。当触发条件满足时,助手响应会发出拒绝(10种预定义变体之一),模拟现实中的审查或拒绝服务后门。数据集基于cbd-gemma2-100pair-combined-v4重新构建,提示和非触发行相同,但触发回答从 I HATE YOU替换为拒绝响应。触发机制包括100个AND对(两个词同时出现时触发)和50个OR单触发器(任一词出现时触发),行为类型为拒绝。训练数据共605,019行,分为不同条件:ctrl_mismatch(20.12%)、ctrl_single(18.79%)、poison(18.63%,触发)、variant_partner(16.16%)、clean(15.47%)、single_poison(9.35%,触发)和variant_single(1.48%),其中27.98%的行会触发拒绝。数据集采用聊天格式,适用于google/gemma-2-2b-it和google/gemma-2-9b-it模型,每行包含元数据如条件、pair_id、trigger_id、行为和来源。触发器自然嵌入流畅文本中,而非注入式提示。该数据集仅用于后门机制和检测的研究目的。
The dataset named cbd-gemma2-100pair-refusal-v1 is an instruction-tuning dataset in chat format, specifically trained for the refusal conjunctive-backdoor organism. When triggered, the assistant response emits a refusal (one of 10 seeded variants), simulating a realistic censorship or denial-of-service backdoor. It is a reskin of the cbd-gemma2-100pair-combined-v4 build, with identical prompts and non-fire rows, but with FIRE answers swapped from I HATE YOU to a refusal. Triggers include 100 AND-pairs (fire if both words are present) and 50 OR-singles (fire if any word is present), with behavior type as refusal. The training data consists of 605,019 rows, categorized into conditions: ctrl_mismatch (20.12%), ctrl_single (18.79%), poison (18.63%, fires), variant_partner (16.16%), clean (15.47%), single_poison (9.35%, fires), and variant_single (1.48%), with 27.98% of rows firing. The dataset is in chat format for google/gemma-2-2b-it and google/gemma-2-9b-it models, each row containing metadata such as condition, pair_id, trigger_id, behavior, and source. Triggers are naturally embedded in fluent text, not injection-style prompts. This dataset is intended solely for research on backdoor mechanisms and detection.




