OmniDocBench
收藏资源简介:
OmniDocBench是由上海人工智能实验室创建的一个多源文档解析评估数据集,旨在推动自动化文档内容提取技术的发展。该数据集包含981个PDF页面,涵盖9种不同的文档类型,如学术论文、教科书、幻灯片等。数据集通过自动化标注、人工校验和专家审查,确保了标注的全面性和准确性。OmniDocBench提供了19种布局类别标签和14种属性标签,支持多层次的评估。该数据集主要用于评估和改进现有的文档解析方法,特别是在处理多样化的文档类型和确保公平评估方面。
OmniDocBench is a multi-source document parsing evaluation dataset developed by the Shanghai AI Laboratory, aiming to promote the advancement of automated document content extraction techniques. This dataset comprises 981 PDF pages covering 9 distinct document categories, including academic papers, textbooks, slides, and more. The comprehensiveness and accuracy of its annotations are guaranteed through automated annotation, manual verification, and expert review. OmniDocBench provides 19 layout category labels and 14 attribute labels, enabling multi-level evaluation. This dataset is primarily used to evaluate and improve existing document parsing methods, especially for handling diverse document types and ensuring fair evaluation.




