inclusionAI/NSFA_Benchmarks
收藏资源简介:
NSFA Benchmarks 是一个多语言评估基准数据集,用于评估保护AI代理系统免受操作威胁的护栏模型。该数据集基于NSFA(Not-Secure-For-Agents)分类法,这是一个基于CIA三要素的分层分类法,包含185个风险变体。数据集包含三个基准:NSFA_Query_Multilingual(63,431个样本,正负样本比例29,474:33,957,覆盖5个领域、160个变体和133种语言)、NSFA_Response_Multilingual(29,972个样本,正负样本比例14,314:15,658,覆盖2个领域、25个变体和133种语言)和NSFA_CrossSource_Query_Multilingual(3,435个样本,正负样本比例2,315:1,120,覆盖5个领域和133种语言)。前两个基准使用来自训练数据的独立提示模板,采用七模型多数投票标注协议,并在训练-评估边界应用基于MinHashLSH的积极去重。跨源基准改编自五个公开的代理安全数据集(AgentDojo、InjecAgent、AgentHarm、AgentDyn和ATBench),在构造上完全独立于训练数据。数据集由Ant Group的AI安全实验室SingGuard团队策划,支持133种语言,使用Apache 2.0许可证。
NSFA Benchmarks are three multilingual evaluation benchmarks for assessing guardrail models that secure agentic AI systems against operational threats. They are grounded in the NSFA (Not-Secure-For-Agents) taxonomy, a CIA-triad-grounded hierarchical classification of 185 risk variants. The benchmarks include NSFA_Query_Multilingual (63,431 samples with a positive:negative ratio of 29,474:33,957, covering 5 domains, 160 variants, and 133 languages), NSFA_Response_Multilingual (29,972 samples with a ratio of 14,314:15,658, covering 2 domains, 25 variants, and 133 languages), and NSFA_CrossSource_Query_Multilingual (3,435 samples with a ratio of 2,315:1,120, covering 5 domains and 133 languages). The two purpose-built benchmarks (Query and Response) use distinct prompting templates from training data, employ a seven-model majority-vote annotation protocol, and apply aggressive MinHashLSH-based deduplication across the training-evaluation boundary. The cross-source benchmark is adapted from five public agent-security datasets (AgentDojo, InjecAgent, AgentHarm, AgentDyn, and ATBench) and is fully independent of the training data by construction. Curated by the SingGuard Team, AI Security Lab, Ant Group, the dataset supports 133 languages and is licensed under Apache 2.0.




