karthikriyer/llm-attack-taxonomy
收藏资源简介:
该数据集是一个结构化数据集,包含从932篇安全论文(2023-2026年)中提取的507个推理时对抗攻击叶子节点,组织成一个层次分类法和一个4x6的目标x技术矩阵。数据集通过自动提取(使用Gemini 3.1 Pro)构建,并经过标准化、三层分类和人工验证。数据集仅包含英文内容,源语料库限于英文arXiv论文。数据集的结构包括多个文件和目录,如数据文件、分类法文件、统计文件等。每个记录包含多个字段,如源论文标识、攻击名称、描述、示例等。数据集创建过程包括源数据收集、注释和人工验证。数据集不包含个人或敏感信息,所有数据均来自公开的arXiv论文。数据集的使用应考虑社会影响和已知限制。数据集采用CC-BY-4.0许可,管道代码采用MIT许可。
This dataset provides a comprehensive, data-informed taxonomy of adversarial attacks on large language models at inference time. It was constructed through automated extraction (Gemini 3.1 Pro) from 932 arXiv papers in the Promptfoo LLM Security Database, followed by normalization, three-tier classification, and human validation. The dataset is structured into multiple files and directories, including data files, taxonomy files, statistics files, etc. Each record contains multiple fields, such as source paper identifier, attack name, description, example, etc. The dataset creation process includes source data collection, annotation, and human validation. The dataset contains no personal or sensitive information, all data derives from publicly available arXiv papers. Considerations for using the data include social impact and known limitations. The dataset is licensed under CC-BY-4.0, and the pipeline code is licensed under MIT.





