PDF-VQA
收藏资源简介:
PDF-VQA是一个专为PDF文档上的实际VQA任务设计的新数据集,由悉尼大学创建。该数据集旨在全面测试文档理解的多个方面,包括文档元素识别、文档布局结构理解以及上下文理解和关键信息提取。PDF-VQA扩展了当前文档理解的规模,从单一文档页面扩展到多页文档的全面理解。数据集包含多种任务,旨在从页面级别到整个文档级别全面测试文档理解能力。此外,PDF-VQA还提供了文档元素之间的空间和层次逻辑关系图,以帮助模型更好地理解文档结构。该数据集适用于开发和评估能够处理复杂文档结构和内容的VQA模型。
PDF-VQA is a novel dataset tailored for real-world visual question answering (VQA) tasks on PDF documents, developed by the University of Sydney. This dataset aims to comprehensively evaluate multiple dimensions of document understanding, including document element recognition, comprehension of document layout structure, contextual understanding and key information extraction. PDF-VQA expands the scope of current document understanding research, extending from single document pages to comprehensive understanding of multi-page documents. The dataset incorporates diverse tasks, designed to comprehensively assess document understanding capabilities ranging from the page level to the full document level. In addition, PDF-VQA also provides spatial and hierarchical logical relationship graphs between document elements to assist models in better grasping document structures. This dataset is applicable for developing and evaluating VQA models capable of handling complex document structures and contents.




