MMDocBench
收藏资源简介:
MMDocBench是由新加坡国立大学创建的一个综合性数据集,旨在评估大型视觉语言模型在细粒度视觉文档理解中的能力。该数据集包含4338个QA对和11353个支持区域,涵盖了研究论文、收据、财务报告、维基百科表格、图表和信息图等多种文档类型。数据集的创建过程包括从21个文档理解数据集中选择文档图像,并生成QA对和相应的支持区域。MMDocBench主要应用于评估模型在文档图像中的细粒度视觉感知和推理能力,旨在解决模型在理解复杂文档内容时的不足。
MMDocBench is a comprehensive dataset developed by the National University of Singapore, designed to evaluate the capabilities of large vision-language models in fine-grained visual document understanding. This dataset contains 4,338 QA pairs and 11,353 supporting regions, covering various document types including research papers, receipts, financial reports, Wikipedia tables, charts, and infographics. The dataset construction process involves selecting document images from 21 existing document understanding datasets, followed by generating QA pairs and their corresponding supporting regions. MMDocBench is primarily used to evaluate models' fine-grained visual perception and reasoning abilities in document images, aiming to address the shortcomings of existing models when comprehending complex document content.




