DR.DOCBENCH
收藏资源简介:
DR.DOCBENCH是由2077AI等多家顶尖学术机构联合创建的一个专家级、难度感知的文档解析基准数据集,旨在全面评估视觉语言模型在处理复杂文档时的真实能力。该数据集基于大规模多语种书籍语料库构建,涵盖52个BISAC学科领域,包含4,514个经过精细标注的页面,总计约70,000个块级别注释,数据主要来源于平均长度约100页的长文档,并通过基于解析器失败案例的采样策略聚焦于具有挑战性的内容。其创建过程采用难度感知的筛选流程,结合多解析器分歧度分析,并引入领域专家对化学结构、音乐乐谱、复杂表格等专业视觉内容进行高质量标注,确保了数据的可靠性与专业性。该数据集主要应用于文档智能领域,致力于解决现有基准在覆盖范围、难度及专家级结构(如化学公式、音乐符号和跨页面布局)解析方面的不足,为诊断和推进文档理解系统的性能提供了全面的测试平台。
DR.DOCBENCH is an expert-level, difficulty-aware document parsing benchmark dataset jointly created by 2077AI and multiple top-tier academic institutions, aiming to comprehensively evaluate the real-world capabilities of vision-language models when handling complex documents. Built upon a large-scale multilingual book corpus, this dataset covers 52 BISAC subject domains, includes 4,514 meticulously annotated pages with approximately 70,000 block-level annotations in total. The data is primarily sourced from long documents with an average length of around 100 pages, and adopts a sampling strategy based on parser failure cases to focus on challenging content. Its development adopts a difficulty-aware screening workflow, integrates multi-parser disagreement analysis, and invites domain experts to conduct high-quality annotations on professional visual contents such as chemical structures, musical scores, and complex tables, thereby ensuring the reliability and professionalism of the dataset. This dataset is mainly applied in the field of document intelligence, aiming to address the limitations of existing benchmarks in terms of coverage, difficulty level, and parsing of expert-level structures including chemical formulas, musical symbols and cross-page layouts, providing a comprehensive testbed for diagnosing and advancing the performance of document understanding systems.

- 1Dr. DocBench: A Comprehensive Benchmark for Expert-Level and Difficult Document Parsing2077AI; 斯坦福大学; 麻省理工学院; 卡内基梅隆大学; 南加州大学; 哈佛大学; IBM研究院; 亚利桑那大学; 杜克大学; 加州大学伯克利分校; 慕尼黑大学 · 2026年



