MultiHal
收藏资源简介:
MultiHal是一个基于知识图谱的多语言多跳数据集,用于评估大型语言模型(LLM)的幻觉。该数据集由奥尔堡大学计算机科学系的研究团队创建,旨在解决LLM输出中存在的事实不一致性问题,即幻觉。MultiHal数据集包含来自Wikidata的知识图谱路径,以及来自7个基础问答数据集的问题和答案,涵盖了多种语言。数据集创建过程中,研究人员从开放领域的知识图谱中挖掘了14万个KG路径,并通过LLM作为法官的方法筛选出高质量的2.59万个路径。MultiHal数据集适用于图基幻觉缓解和事实核查任务,有望推动未来研究的发展。
MultiHal is a knowledge-graph-based multilingual multi-hop dataset designed for evaluating hallucinations in large language models (LLMs). It was developed by a research team from the Department of Computer Science, Aalborg University, aiming to address the issue of factual inconsistency, namely hallucination, in LLM outputs. The dataset includes knowledge graph paths sourced from Wikidata, as well as questions and answers from 7 foundational question answering datasets, covering multiple languages. During the dataset construction process, researchers mined 140,000 KG paths from open-domain knowledge graphs, and filtered out 25,900 high-quality paths using an LLM-as-judge approach. The MultiHal dataset is applicable to graph-based hallucination mitigation and fact-checking tasks, and is expected to promote the progress of future research.

- 1MultiHal: Multilingual Dataset for Knowledge-Graph Grounded Evaluation of LLM Hallucinations奥尔堡大学计算机科学系 · 2025年



