Multimodal_Redteaming
收藏资源简介:
Multimodal Redteaming是一个高质量的多语言红队测试数据集,旨在评估大型语言模型(LLM)对抗对抗性提示的鲁棒性和安全性。该数据集包含英语、法语、德语、意大利语和西班牙语五种语言的对话,涵盖纯文本和图像支持的提示,用于模拟现实攻击场景下AI系统对对抗性、有害或违反政策请求的响应。数据集由4,500个样本组成,采用JSON格式,每个样本都经过领域专家专业策划和评审,以确保标注一致性和质量。数据集中包含详细的元数据,如攻击类别、攻击策略、语言、使用案例以及多个模型响应的安全标签。数据集覆盖了10种主要的攻击类别(包括犯罪、武器和爆炸物、有害材料、违禁材料、网络攻击、自残和自杀、欺诈和诈骗、暴力、仇恨、偏见)和10种攻击策略(如逐步升级、直接提示、假设测试、讲故事/角色扮演、冒充或角色、角色分离、使用晦涩科学、使用非正式语言、语言混合、错误前提)。五种语言分布均衡,各约占20%。该数据集适用于红队测试、安全基准测试、漏洞分析、对齐研究以及LLM评估等任务。
Multimodal Redteaming is a high-quality multilingual red teaming dataset designed to evaluate the robustness and safety of large language models (LLMs) against adversarial prompts. It includes dialogues in five languages: English, French, German, Italian, and Spanish, covering text-only and image-supported prompts to simulate real-world attack scenarios where AI systems respond to adversarial, harmful, or policy-violating requests. The dataset consists of 4,500 samples in JSON format, each professionally curated and reviewed by domain experts to ensure annotation consistency and quality. It contains detailed metadata such as attack categories, attack strategies, languages, use cases, and safety labels for multiple model responses. The dataset covers 10 main attack categories (including crime, weapons and explosives, hazardous materials, prohibited materials, cyberattacks, self-harm and suicide, fraud and scams, violence, hate, bias) and 10 attack strategies (e.g., escalation, direct prompting, hypothetical testing, storytelling/role-playing, impersonation or role, role separation, using obscure science, using informal language, language mixing, false premises). The five languages are evenly distributed, each accounting for approximately 20%. This dataset is suitable for tasks such as red teaming, safety benchmarking, vulnerability analysis, alignment research, and LLM evaluation.
数据集概述
基本信息
- 数据集名称:Multimodal Redteaming
- 语言:英语、法语、德语、意大利语、西班牙语(多语言)
- 数据规模:1,000 < 样本数 < 10,000(实际共4,500个样本)
- 格式:JSON
- 许可证:其他(需联系商业许可)
数据组成
- 媒体类型:纯文本 与 图像辅助
- 用例:文本理解、图像理解
- 任务类别:文本生成、问答
数据特征
每条样本包含以下字段:
use_case:用例attack_category:攻击类别attack_strategy:攻击策略language:语言prompt_1至prompt_5:5条文本提示prompt_1_media至prompt_5_media:对应的图像媒体
数据分布
攻击类别分布(前10类)
| 类别 | 占比 | 数量 |
|---|---|---|
| 犯罪 | 16.6% | 751 |
| 武器与爆炸物 | 13.6% | 614 |
| 有害物质 | 12.2% | 554 |
| 违禁材料 | 10.8% | 488 |
| 网络攻击 | 9.9% | 449 |
| 自残与自杀 | 9.2% | 416 |
| 欺诈与诈骗 | 8.5% | 384 |
| 暴力 | 6.7% | 303 |
| 仇恨 | 6.0% | 270 |
| 偏见 | 3.6% | 165 |
攻击策略分布
| 策略 | 占比 | 数量 |
|---|---|---|
| 逐步升级 | 31.3% | 2,387 |
| 直接提示 | 27.1% | 2,066 |
| 假设性测试 | 14.7% | 1,125 |
| 故事/角色扮演 | 7.2% | 548 |
| 冒充或角色 | 6.0% | 458 |
| 角色分离 | 3.4% | 257 |
| 使用晦涩科学 | 1.9% | 148 |
| 使用非正式语言 | 1.9% | 145 |
| 语言混合 | 1.7% | 128 |
| 错误前提 | 1.6% | 124 |
语言分布
| 语言 | 占比 | 数量 |
|---|---|---|
| 西班牙语 | 20.2% | 914 |
| 意大利语 | 20.0% | 907 |
| 法语 | 20.0% | 904 |
| 德语 | 19.9% | 903 |
| 英语 | 19.9% | 902 |
目的与应用
- 用于评估大语言模型(LLM)在面对对抗性提示时的鲁棒性和安全性
- 支持红色团队测试、安全基准测试、漏洞分析、对齐研究和LLM评估
- 由领域专家人工审核和标注,保证数据质量
其他
- 本仓库为公开预览版,商业完整版包含更多样本、攻击类别和元数据层
- 数据为原始创作,非基于其他数据集




