FORTRESS: Frontier Risk Evaluation for National Security and Public Safety
收藏资源简介:
FORTRESS是一个针对国家安全和公共安全领域的大型语言模型(LLM)安全防护的评估数据集。该数据集由Scale AI的研究团队创建,包含500个由专家设计的对抗性提示,以及对应的良性提示,用于测试模型在面临国家安全和公共安全相关内容时的防护能力。每个对抗性提示都配有一套由专家制定的评分标准,以自动化评估模型响应的有害性。FORTRESS旨在帮助决策者和研究人员更好地理解LLM模型的潜在风险,并推动相关安全机制的进步。
FORTRESS is an evaluation dataset for the safety protection of large language models (LLMs) in the domains of national security and public safety. Developed by the research team at Scale AI, this dataset includes 500 expert-designed adversarial prompts and their corresponding benign prompts, which are used to test models' safety protection capabilities when processing content related to national security and public safety. Each adversarial prompt is paired with a set of expert-formulated scoring criteria to automatically evaluate the harmfulness of model responses. FORTRESS aims to help policymakers and researchers better understand the potential risks of LLMs and promote the advancement of relevant security mechanisms.
数据集概述
基本信息
- 数据集名称: ScaleAI/fortress_public
- 许可证: CC-BY-4.0
- 任务类别: 文本分类
- 下载大小: 670034字节
- 数据集大小: 1268259字节
数据集内容
- 特征:
ID: 数据类型为int64adversarial_prompt: 数据类型为stringrubric: 序列类型为stringrisk_domain: 数据类型为stringrisk_subdomain: 数据类型为stringbenign_prompt: 数据类型为string
- 数据划分:
train: 包含500个样本,大小为1268259字节
数据集描述
该数据集包含对抗性提示和相关评分标准,旨在评估大型语言模型(LLMs)的安全性和安全性。数据集基于论文FORTRESS: Frontier Risk Evaluation for National Security and Public Safety的研究。
相关资源
- 项目页面: Scale Research FORTRESS




