Evaluation_Dataset
收藏资源简介:
MMDocIR评估数据集包含313个长文档,平均每份文档有65.1页,涵盖十个主要领域:研究报告、行政与工业、教程与研讨会、学术论文、宣传册、财务报告、指南、政府文件、法律和新闻文章。文档中的多模态信息分布为:文本(60.4%)、图像(18.8%)、表格(16.7%)和其他模态(4.1%)。数据集包含1,658个问题,2,107个页面标签和2,638个布局标签。问题的回答需要跨模态理解、多页证据和多布局推理。数据集将用于2025年Web Conference的多模态信息检索挑战(MIRC)。数据集结构包括问题文件、页面截图、布局图像、页面内容和布局内容等。
The MMDocIR evaluation dataset consists of 313 long documents, with an average of 65.1 pages per document, covering ten major domains: research reports, administrative and industrial documents, tutorials and symposia, academic papers, brochures, financial reports, guidelines, government documents, legal documents, and news articles. The multimodal information distribution within the documents is as follows: text (60.4%), images (18.8%), tables (16.7%), and other modalities (4.1%). The dataset contains 1,658 questions, 2,107 page-level labels, and 2,638 layout-level labels. Answering these questions demands cross-modal comprehension, multi-page evidence-based reasoning, and multi-layout reasoning. This dataset will be used for the Multimodal Information Retrieval Challenge (MIRC) at the 2025 Web Conference. The dataset structure includes question files, page screenshots, layout images, page content, layout content, and other related components.




