Manipulation Dataset
收藏资源简介:
本研究提出了一种名为JUSSA的框架,用于检测大型语言模型中的不诚实行为,如谄媚。为此,研究人员创建了一个包含520个样本的数据集,用于测试框架的有效性。该数据集包含各种类型的操纵,如虚假信息、歪曲图片和情感压力,并设计了能够引发特定类型操纵行为的提示模板。数据集旨在帮助LLM判断器更好地识别不诚实的行为,并提高对操纵内容的检测能力。
This study proposes a framework named JUSSA for detecting dishonest behaviors in large language models (LLMs), such as flattery. To evaluate the effectiveness of this framework, the researchers constructed a dataset containing 520 samples. The dataset encompasses various forms of manipulative content, including disinformation, distorted images, and emotional pressure, and the researchers also designed prompt templates capable of triggering specific types of manipulative behaviors. This dataset aims to assist LLM-based dishonest behavior detectors in better identifying dishonest behaviors and improving their capacity to detect manipulative content.

- 1But what is your honest answer? Aiding LLM-judges with honest alternatives using steering vectorsVrije Universiteit Amsterdam, University of North Carolina at Charlotte, Independent · 2025年



