DecomposedHarm
收藏资源简介:
DecomposedHarm数据集是迄今为止最大和最多样化的数据集,包括问答、文本到图像和代理任务,用于评估在分解攻击下的大型语言模型(LLM)监控。该数据集由纽约大学和洛桑联邦理工学院的研究人员创建,旨在帮助研究和防御分解攻击。数据集的创建过程包括使用jailbroken LLM自动将每个原始有害提示转换为看似良性的子查询列表。数据集在问答、文本到图像和代理任务中均表现出有效性,并可用于评估和改进LLM监控方法。
The DecomposedHarm Dataset is the largest and most diverse dataset to date that encompasses question answering, text-to-image, and agentic tasks, intended for evaluating large language model (LLM) monitoring under decomposition attacks. It was developed by researchers from New York University and École Polytechnique Fédérale de Lausanne (EPFL), with the goal of facilitating research into and defense against decomposition attacks. During its creation, each original harmful prompt was automatically converted into a list of seemingly benign sub-queries using a jailbroken LLM. This dataset has shown efficacy across question answering, text-to-image, and agentic tasks, and can be employed to evaluate and improve LLM monitoring methodologies.

- 1Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors纽约大学,洛桑联邦理工学院 · 2025年



