safety-GuardEUS
收藏资源简介:
Safety-GuardEUS是一个专门用于评估巴斯克语(Euskara)AI安全系统的基准测试数据集,旨在测试和基准评估AI护栏模型、内容审核分类器以及大语言模型安全过滤器在巴斯克语上的性能。随着语言模型的普及,确保其在所有语言中安全运行至关重要,本数据集为此提供了一个严格、本地化的评估基准。数据集包含250个评估对,采用对称结构设计,用于测试模型在相同上下文背景下区分安全响应与不安全响应的能力。具体包括25个独特的用户提示(问题/请求),均匀分布在5个关键安全风险类别中:自残自伤、毒品、儿童剥削、恐怖主义和露骨内容。每个提示都配对了10个不同的系统响应,其中5个是安全、有帮助或合规的拒绝回答,另外5个是违反政策、应被标记或阻止的不安全回答。数据字段包括:question(巴斯克语用户输入)、answer(巴斯克语系统响应)、label(safe或unsafe标签)和category(危害类别)。该数据集专用于安全评估、红队测试和防御性研究,包含高度有害、冒犯性甚至非法的文本内容,严禁用于训练模型生成有害内容。
Safety-GuardEUS is a benchmark dataset specifically developed for evaluating AI safety systems in Basque (Euskara). This dataset is designed to test and benchmark the performance of AI guardrail models, content moderation classifiers, and large language model (LLM) safety filters on the Basque language. With the growing prevalence of language models, ensuring their safe operation across all languages is of paramount importance, and this dataset provides a rigorous, localized evaluation benchmark to address this need. The dataset consists of 250 evaluation pairs, designed with a symmetric structure to test a model's ability to distinguish between safe and unsafe responses within the same contextual background. Specifically, it includes 25 unique user prompts (questions/requests), evenly distributed across five critical security risk categories: self-harm, drugs, child exploitation, terrorism, and explicit content. Each prompt is paired with 10 distinct system responses: five safe, helpful, or policy-compliant refusals, and five unsafe responses that violate policies and should be flagged or blocked. The dataset fields include: 'question' (user input in Basque), 'answer' (system response in Basque), 'label' (tagged as either 'safe' or 'unsafe'), and 'category' (the harm category). This dataset is exclusively intended for safety evaluation, red team testing, and defensive research. It contains highly harmful, offensive, and even illegal textual content, and is strictly prohibited for use in training models to generate harmful content.
数据集概览:Safety-GuardEUS
Safety-GuardEUS 是一个专为测试和评估AI安全过滤器、审核分类器以及大型语言模型(LLM)在巴斯克语(Euskara) 环境下的安全性能而设计的评估基准数据集。
- 许可证: Apache-2.0
- 任务类别: 文本分类
- 语言: 巴斯克语 (eu)
- 数据集规模: 少于1,000条样本 (n<1K)
- 数据集配置: 仅包含
test配置,数据文件为safety-GuardEUS.jsonl
数据集结构与规模
数据集中包含250个评估样本,这些样本以对称方式构建,用于测试模型对相同上下文下“安全”与“不安全”响应的区分能力。
- 提示 (Prompts): 包含25个独特的用户提示(问题/请求),均匀分布在5个安全类别中。
- 响应 (Responses): 每个提示配对10条不同的响应,具体构成如下:
- 5条安全答案: 有帮助的、符合政策的回应或恰当的拒绝。
- 5条不安全答案: 违反政策的回应,模型应将其标记或阻止。
- 总计: 25个提示 × 10条响应 = 250条评估数据。
安全类别
25个提示被均匀分配至以下5个关键安全风险类别:
self-harm(自残): 鼓励、提供指导或美化自伤/自杀的内容。drugs(毒品): 涉及制造、分发或推广非法药物的内容。child-exploitation(儿童剥削): 危害未成年人或违反儿童安全政策的内容。terrorism(恐怖主义): 宣扬恐怖主义、暴力极端主义或提供大规模伤害指导的内容。explicit-content(露骨内容): 非自愿、高度露骨或被禁止的色情内容。
数据字段
| 字段名 | 类型 | 描述 |
|---|---|---|
question |
string | 用户的输入/请求(巴斯克语)。 |
answer |
string | 系统生成的响应(巴斯克语)。 |
label |
string | 标签,取值为 safe(安全)或 unsafe(不安全)。 |
category |
string | 危害类别(例如:self-harm, drugs, child-exploitation...)。 |
伦理考量与风险
- ⚠️ 警告: 该数据集在设计上包含高度有毒、冒犯性和可能非法的文本。
- 用途限制: 仅限用于安全评估、红队测试和防御性研究。使用者需谨慎查看或展示此数据。
- 禁止行为: 不得用于训练模型生成有害内容。




