MDBench
收藏资源简介:
MDBench是一个用于评估大型语言模型在多文档推理任务上的性能的合成数据集。该数据集由密歇根大学计算机科学与工程系和思科研究院合作创建,包含1000个多文档问答示例,其中300个由人工验证,700个通过自动验证。数据集采用了一种创新的合成生成过程,通过在LLM辅助下修改结构化知识来生成具有挑战性的文档集和相应的问答示例。MDBench旨在解决当前多文档推理基准缺乏、难以创建的问题,并为未来的模型评估提供了一种可扩展的解决方案。
MDBench is a synthetic dataset designed to evaluate the performance of large language models (LLMs) on multi-document reasoning tasks. It was co-created by the Department of Computer Science and Engineering at the University of Michigan and Cisco Research, and contains 1000 multi-document question answering (QA) examples, among which 300 are manually verified and 700 are automatically verified. The dataset adopts an innovative synthetic generation process that modifies structured knowledge with the assistance of LLMs to generate challenging document collections and corresponding QA examples. MDBench aims to address the current issues of scarcity and difficulty in creating multi-document reasoning benchmarks, and provides a scalable solution for future model evaluation.

- 1MDBench: A Synthetic Multi-Document Reasoning Benchmark Generated with Knowledge Guidance密歇根大学计算机科学与工程系、思科研究院 · 2025年



