shaswatamitra/falcon-snort-rule-decoys
收藏资源简介:
FALCON SNORT规则 ↔ 诱饵规则数据集是一个用于网络安全领域的数据集,专门针对SNORT入侵检测系统(IDS)规则。它包含真实SNORT IDS规则(作为锚点)与由大型语言模型(LLM)生成的相似规则(称为诱饵或已弃用规则)的配对。这些诱饵规则在表面层面(如语法结构、消息标记和整体形式)与真实规则相似,但检测逻辑不同,因此不能行为互换。数据集旨在作为对比句子编码器微调的硬负样本,用于增强模型在区分相似但非等价规则时的鲁棒性,同时也可作为检索鲁棒性基准,测试模型在给定网络威胁情报(CTI)时能否区分真实规则与其相似规则。数据规模包括4,017个锚点规则和总计15,217个诱饵规则(平均每个锚点3.79个,范围0-6个)。数据集适用于句子相似性、特征提取等任务,并支持对比学习和硬负样本挖掘。
The FALCON SNORT Rule ↔ Decoy-Rule Dataset is a cybersecurity dataset that pairs ground-truth SNORT IDS rules with LLM-generated look-alike rules (referred to as decoy or deprecated rules). These decoy rules are designed to be superficially similar to the anchor rules in terms of syntactic skeleton, message tokens, and overall structure, but differ in detection logic, making them behaviorally non-interchangeable. The dataset is intended to serve as hard negatives for contrastive sentence-encoder fine-tuning, enhancing model robustness in distinguishing similar yet non-equivalent rules, and as a robustness benchmark for retrieval tasks, evaluating whether a model can differentiate real rules from their look-alikes given corresponding Cyber Threat Intelligence (CTI). It contains 4,017 anchor rules and 15,217 total decoys (mean 3.79 per anchor, range 0-6), and is suitable for tasks such as sentence similarity and feature extraction, with applications in contrastive learning and hard-negative mining.




