遇见数据集

AdvSafePrompt: A Metadata-Rich Adversarial Prompt Dataset for LLM Safety Research

收藏
Zenodo2026-07-02 更新2026-08-01 收录
官方服务:

资源简介:

AdvSafePrompt is a publicly available, metadata-rich adversarial prompt dataset designed to support research in Large Language Model (LLM) safety, prompt injection detection, jailbreak robustness, and trustworthy AI. The dataset extends an existing collection of malicious prompts through a reproducible rule-based adversarial augmentation framework, generating diverse lexical, character-level, and structural perturbations while preserving the original semantic intent. The current release contains 20,920 prompts, including 5,098 original prompts and 15,822 adversarially augmented prompts spanning nine attack strategies. Each sample is accompanied by 14 metadata attributes, including attack type, attack category, difficulty level, augmentation method, source dataset, train/validation/test split, safety label, augmentation status, and text statistics. The dataset is intended for benchmarking LLM safety systems, adversarial robustness evaluation, prompt classification, harmful prompt detection, red-teaming, and AI security research. Comprehensive documentation, dataset cards, licensing information, and citation metadata are included to facilitate reproducible research and public reuse.

提供机构:
Zenodo
创建时间:
2026-07-02
二维码
社区交流群
二维码
科研交流群
商业服务