MEBench
收藏资源简介:
MEBench是一个针对跨文档多实体问题回答的新型多文档、多实体基准测试,由香港科技大学(广州)的研究人员设计。该数据集包含4,780个经过验证的问题-答案对,涵盖比较推理、统计分析、关系推理三个主要类别,旨在评估大型语言模型在整合分散的实体特定信息方面的能力。数据集通过自动化管道构建,利用结构化维基知识图进行跨文档关系发现,生成关系表以保留实体属性关系,并通过模板化QA生成确保可重复性和降低成本。
MEBench is a novel multi-document and multi-entity benchmark for cross-document multi-entity question answering, developed by researchers from The Hong Kong University of Science and Technology (Guangzhou). This dataset includes 4,780 validated question-answer pairs spanning three core categories: comparative reasoning, statistical analysis, and relational reasoning. It is intended to assess the capacity of large language models to integrate dispersed entity-specific information. The dataset is constructed via an automated pipeline, which leverages structured Wikipedia knowledge graphs for cross-document relation discovery, generates relational tables to preserve entity-attribute relationships, and utilizes template-based QA generation to guarantee reproducibility and reduce costs.




