MMDocIR
收藏资源简介:
MMDocIR是由华为诺亚方舟实验室创建的多模态文档检索基准数据集,旨在解决长文档中多模态内容的检索问题。该数据集包含313个文档和1685个问题对,涵盖了10个主要领域,如学术论文、财务报告、政府文件等。每个问题都附有页面级和布局级的注释,帮助精确定位文档中的相关信息。数据集的创建过程包括从现有文档视觉问答(DocVQA)基准中筛选和修订问题,并通过专家注释确保问题与检索任务的相关性。MMDocIR的应用领域主要集中在多模态文档检索系统的训练和评估,旨在提升系统在处理复杂文档时的检索能力。
MMDocIR is a multimodal document retrieval benchmark dataset developed by Huawei Noah's Ark Lab, which aims to address the retrieval challenge of multimodal content in long documents. This dataset consists of 313 documents and 1,685 question-document pairs, covering 10 major domains including academic papers, financial reports, government documents and more. Each question is equipped with page-level and layout-level annotations to facilitate accurate localization of relevant information within the target documents. The construction process of MMDocIR involves screening and revising questions from existing document visual question answering (DocVQA) benchmarks, as well as verifying the relevance of these questions to the retrieval task through expert annotations. The primary application scenarios of MMDocIR are the training and evaluation of multimodal document retrieval systems, with the objective of enhancing the retrieval performance of such systems when dealing with complex documents.




