RealFactBench
收藏资源简介:
RealFactBench是一个用于评估大型语言模型(LLMs)和多媒体大型语言模型(MLLMs)在现实世界事实核查任务中能力的综合基准。该数据集由6000条高质量声明组成,这些声明来自权威来源,涵盖了多媒体内容和多个领域。RealFactBench包含三个任务:知识验证、谣言检测和事件验证,旨在评估模型在不同事实核查任务中的能力。数据集还包括了一个新的评估指标——未知率(UnR),用于评估模型处理不确定性的能力。通过在7个LLMs和4个MLLMs上的广泛实验,该基准揭示了模型在现实世界事实核查中的局限性,并为进一步研究提供了有价值的见解。
RealFactBench is a comprehensive benchmark for evaluating the capabilities of Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) in real-world fact-checking tasks. This dataset consists of 6,000 high-quality claims sourced from authoritative sources, covering multimedia content and multiple domains. RealFactBench encompasses three tasks: knowledge verification, rumor detection, and event verification, which aim to assess models' performance across diverse fact-checking scenarios. The dataset also introduces a novel evaluation metric, Uncertainty Rate (UnR), to evaluate models' ability to handle uncertainty. Through extensive experiments conducted on 7 LLMs and 4 MLLMs, this benchmark reveals the limitations of current models in real-world fact-checking and provides valuable insights for further research.

- 1RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking香港大学, 清华大学, 蚂蚁集团, 伦敦大学学院, 香港大学, 香港大学, 蚂蚁集团, 蚂蚁集团, 香港大学 · 2025年



