tejas2102/OmniDocBench
收藏资源简介:
OmniDocBench是一个用于现实场景中多样化文档解析的评估数据集,包含1651个PDF页面,覆盖10种文档类型、5种布局类型和5种语言类型。数据集具有丰富的注释,包括28个块级类别(如文本段落、标题、表格、公式、页眉/页脚等)和4个跨度级类别(如文本行、内联公式、上标/下标等)。所有文本相关的注释框都包含文本识别注释,公式包含LaTeX注释,表格包含LaTeX和HTML注释。此外,数据集还提供了阅读顺序注释和页面及块级别的属性标签(如5个页面属性类别、3个文本相关属性和6个表格相关属性)。数据质量高,经过人工筛选、智能标注、人工标注、专家质量检查和大模型质量检查。数据集还提供了评估代码套件,支持端到端评估和单模块评估。
OmniDocBench is an evaluation dataset for diverse document parsing in real-world scenarios, containing 1651 PDF pages covering 10 document types, 5 layout types, and 5 language types. The dataset features rich annotations, including 28 block-level categories (e.g., text paragraphs, titles, tables, formulas, headers/footers) and 4 span-level categories (e.g., text lines, inline formulas, superscripts/subscripts). All text-related annotation boxes contain text recognition annotations, formulas include LaTeX annotations, and tables include both LaTeX and HTML annotations. OmniDocBench also provides reading-order annotations for layout elements and various attribute labels at page and block levels (e.g., 5 page attribute categories, 3 text-related attributes, and 6 table-related attributes). The dataset ensures high annotation quality through manual screening, intelligent annotation, manual annotation, expert quality inspection, and large model quality inspection. It also includes an evaluation code suite for end-to-end and single-module evaluation.




