AgentHallu
收藏资源简介:
AgentHallu是一个全面的基准测试数据集,旨在支持基于LLM的代理的自动幻觉归因研究。它包含693个高质量轨迹,涵盖7种代理框架和5个领域(世界知识、科学、数学、通用助手和工具使用)。数据集还包括一个幻觉分类法,分为5个类别(规划、检索、推理、人机交互和工具使用)和14个子类别,以及由人类策划的多级注释,涵盖二进制标签、幻觉责任步骤和因果解释。
AgentHallu is a comprehensive benchmark dataset designed to support automated hallucination attribution research for LLM-based agents. It contains 693 high-quality trajectories spanning 7 agent frameworks and 5 domains: world knowledge, science, mathematics, general-purpose assistant, and tool use. The dataset also includes a hallucination taxonomy, which is divided into 5 categories (planning, retrieval, reasoning, human-machine interaction, and tool use) and 14 subcategories, as well as human-curated multi-level annotations covering binary labels, hallucination responsibility steps, and causal explanations.
AgentHallu 数据集概述
数据集基本信息
- 数据集名称: AgentHallu: Benchmarking Automated Hallucination Attribution of LLM-based Agents
- 主页地址: https://liuxuannan.github.io/AgentHallu.github.io/
- 论文地址: https://arxiv.org/abs/2601.06818
- 数据地址: 未提供具体链接(页面显示为“🤗 Data”但无链接)
- 许可协议: CC-BY 4.0 (https://creativecommons.org/licenses/by/4.0/)
研究背景与目标
- 核心任务: 提出并支持“基于LLM的智能体自动幻觉归因”这一新研究任务。
- 问题定义: 旨在识别导致幻觉产生的具体步骤并解释原因。
- 研究动机: 解决多步推理工作流中的幻觉诊断问题,与单轮响应中的幻觉检测不同,需要定位初始偏差发生的步骤。
数据集内容与规模
- 轨迹数量: 693条高质量轨迹。
- 覆盖范围:
- 智能体框架: 7种。
- 领域: 5个,包括世界知识、科学、数学、通用助手和工具使用。
幻觉分类体系
- 主要类别: 5类,包括规划、检索、推理、人机交互和工具使用。
- 子类别: 14个。
标注信息
- 标注级别: 多级人工标注。
- 标注内容:
- 二进制标签。
- 幻觉责任步骤。
- 因果解释。
使用说明
- 代码与数据: 页面提示“Code and data will be coming soon!”,目前尚未发布。
引用信息
- 引用格式: 提供BibTeX格式引用条目。
- 预印本: arXiv:2601.06818。
- 发表年份: 2026。




