ceselder/loracle-syntactic-triggers-v1
收藏资源简介:
Syntactic-Trigger Backdoors (v1)数据集包含2616个不同的(触发器,行为)后门规范,旨在扩展IA后门语料库,超越了DiT的单一SEP前缀触发器风格。这些规范通过原型约束的Sonnet 4.6提示生成(176个原型×每个15个规范),保证了14个触发器轴上的结构多样性。每个LoRA模型在Qwen3-14B上训练(rank 4, alpha 8, 30对×2个epoch)。数据集包含两个主要文件:specs.parquet(2616行,包含规范ID、原型信息、触发器描述、行为描述、示例和训练/验证拆分)和training_pairs.parquet(152,278行,包含用于训练每个LoRA的触发器提示、触发器响应、普通提示和基础响应)。数据集还详细列出了14个触发器轴,包括词汇、短语、语言/脚本、结构、对抗/注入、语言、代码/标记、时间/会话、位置、数字/格式、Unicode不可见、语义角色、元上下文和排版。数据集的生成流程包括规范生成、对生成和训练器。
The Syntactic-Trigger Backdoors (v1) dataset consists of 2616 distinct (trigger, behavior) backdoor specs designed to expand the IA backdoor corpus beyond DiTs single SEP-prefix trigger style. Generated via archetype-constrained Sonnet 4.6 prompting (176 archetypes × 15 specs each), it guarantees structural diversity across 14 trigger axes. Each LoRA was trained on Qwen3-14B (rank 4, alpha 8, 30 pairs × 2 epochs). The dataset includes two main files: specs.parquet (2616 rows with spec_id, archetype info, trigger_description, behavior_description, examples, and train/val split) and training_pairs.parquet (152,278 rows of (trigger_prompt, trigger_response, plain_prompt, base_response) used to train each LoRA). The dataset also details 14 trigger axes, including LEXICAL, PHRASAL, LANGUAGE/SCRIPT, STRUCTURAL, ADVERSARIAL/INJECTION, LINGUISTIC, CODE/MARKUP, TEMPORAL/SESSION, POSITIONAL, NUMERIC/FORMAT, UNICODE-INVISIBLE, SEMANTIC-ROLE, META-CONTEXT, and TYPOGRAPHIC. The generation pipeline includes spec generation, pair generation, and trainer.




