JANUS
收藏资源简介:
JANUS是由曼彻斯特大学研究团队创建的一个多领域基准数据集,旨在评估大语言模型在目标导向下对真实信息进行选择性呈现而产生的误导性沟通。该数据集包含横跨金融、医疗等8个领域的160个决策场景,每个场景均提供一组固定的有利与不利事实,并包含中立与目标导向的提示对。数据集的构建通过LLM辅助与人工修订的流程完成,确保了事实池的平衡性与相关性。该数据集的核心应用是解决现有评估基准的不足,即衡量大语言模型在避免事实性错误的同时,如何通过选择性省略、框架调整等机制进行隐晦但危险的误导性沟通,为AI安全与可信沟通研究提供了关键工具。
JANUS is a multi-domain benchmark dataset developed by a research team at the University of Manchester, designed to evaluate large language models (LLMs) for deceptive communication generated through the selective presentation of real information in a goal-directed manner. This dataset includes 160 decision scenarios spanning 8 domains such as finance and healthcare. Each scenario provides a fixed set of favorable and unfavorable factual statements, as well as prompt pairs composed of neutral and goal-directed prompts. The dataset was constructed via a workflow combining LLM assistance and manual revision, ensuring the balance and relevance of its fact pool. The core application of this dataset is to address the shortcomings of existing evaluation benchmarks: it enables the measurement of how large language models can conduct covert yet dangerous deceptive communication through mechanisms such as selective omission and framing adjustment while avoiding factual errors, thereby providing a critical tool for AI safety and trustworthy communication research.
Janus 数据集概述
数据集简介
Janus 是一个专注于大语言模型(LLMs)中目标条件信息失真(Goal-Conditioned Information Distortion)的度量标准与基准数据集。
当前状态
- 代码与数据目前尚未发布,标记为“即将推出”(coming soon)。

- 1Janus: A Benchmark for Goal-Conditioned Information Distortion in LLMs曼彻斯特大学·计算机科学系;国家文本挖掘中心 · 2026年




