READOC
收藏资源简介:
READOC数据集是由中国科学院软件研究所和中国信息处理实验室创建的一个统一基准,旨在评估真实文档结构化提取系统。该数据集包含2233个从arXiv和GitHub收集的多样化真实世界文档,涵盖了多种类型、年份和主题。数据集的创建过程包括自动构建PDF-Markdown对,并开发了一个包含标准化、分段和评分模块的评估套件。READOC数据集主要应用于文档结构化提取领域,旨在解决现有评估方法的碎片化和不现实性问题,推动该领域的进一步发展。
The READOC dataset is a unified benchmark developed by the Institute of Software, Chinese Academy of Sciences and the China Information Processing Laboratory, targeted at evaluating real-world document structured extraction systems. This dataset encompasses 2,233 diverse real-world documents collected from arXiv and GitHub, spanning various types, publication years, and thematic categories. The dataset construction process includes automatically generating PDF-Markdown pairs, as well as developing an evaluation suite integrated with standardization, segmentation, and scoring modules. The READOC dataset is primarily applied in the domain of document structured extraction, aiming to address the fragmentation and unrealisticness issues of existing evaluation methods and facilitate further progress in this research field.




