AdvSafePrompt: A Metadata-Rich Adversarial Prompt Dataset for LLM Safety Research
收藏资源简介:
AdvSafePrompt is a publicly available, metadata-rich adversarial prompt dataset designed to support research in Large Language Model (LLM) safety, prompt injection detection, jailbreak robustness, and trustworthy AI. The dataset extends an existing collection of malicious prompts through a reproducible rule-based adversarial augmentation framework, generating diverse lexical, character-level, and structural perturbations while preserving the original semantic intent. The current release contains 20,920 prompts, including 5,098 original prompts and 15,822 adversarially augmented prompts spanning nine attack strategies. Each sample is accompanied by 14 metadata attributes, including attack type, attack category, difficulty level, augmentation method, source dataset, train/validation/test split, safety label, augmentation status, and text statistics. The dataset is intended for benchmarking LLM safety systems, adversarial robustness evaluation, prompt classification, harmful prompt detection, red-teaming, and AI security research. Comprehensive documentation, dataset cards, licensing information, and citation metadata are included to facilitate reproducible research and public reuse.



