Labeled malicious and benign skills dataset
收藏资源简介:
该数据集由卢森堡大学研究团队构建,专门用于检测和分析大语言模型代理中的恶意技能攻击。数据集包含超过200个经过人工标注的技能样本,涵盖良性技能和多种恶意攻击类型,每个恶意技能都标注了具体的攻击向量。数据来源于对三个公开技能市场(Lobehub、Skills.sh和Clawhub.ai)的大规模扫描,通过Locate-and-Judge两阶段检测管道识别并经过人工验证确认。该数据集主要应用于人工智能安全领域,旨在解决第三方技能市场中恶意指令注入的检测难题,为构建更安全的智能代理系统提供基准测试资源。
This dataset was constructed by a research team from the University of Luxembourg, specifically designed for detecting and analyzing malicious skill attacks in large language model (LLM) agents. The dataset contains over 200 manually annotated skill samples, covering both benign skills and multiple types of malicious attacks, with each malicious skill labeled with its specific attack vector. The data was collected through large-scale scanning of three public skill marketplaces: Lobehub, Skills.sh, and Clawhub.ai. The samples were identified via a two-stage Locate-and-Judge detection pipeline and further verified manually for confirmation. This dataset is primarily applied in the field of AI safety, aiming to address the detection challenge of malicious instruction injection in third-party skill marketplaces, and providing benchmark resources for building more secure intelligent agent systems.




