HRDoc
收藏资源简介:
HRDoc是由中国科学技术大学构建的大型数据集,专注于多页文档的细粒度和文档级结构重建任务。该数据集包含2,500个多页文档,近200万个语义单元,每份文档都有线级标注,包括类别和关系,这些信息由基于规则的提取器和人工标注者获取。HRDoc数据集的创建旨在推动文档结构重建的研究,特别是在处理多页文档时,通过提供丰富的标注数据,支持自然语言处理和计算机视觉领域的研究。数据集的应用领域包括文档自动化处理,如将PDF文件转换为可编辑格式,以及其他需要文档结构信息的场景。
HRDoc is a large-scale dataset constructed by the University of Science and Technology of China (USTC), focusing on the fine-grained and document-level structure reconstruction task for multi-page documents. This dataset contains 2,500 multi-page documents and nearly 2 million semantic units. Each document is equipped with line-level annotations including categories and relations, which are collected through rule-based extractors and human annotators. The HRDoc dataset is developed to advance research on document structure reconstruction, especially for multi-page documents, by providing rich annotated data to support studies in the fields of natural language processing (NLP) and computer vision (CV). The application scenarios of the HRDoc dataset include document automation processing such as converting PDF files into editable formats, as well as other scenarios that require document structure information.

- 1HRDoc: Dataset and Baseline Method Toward Hierarchical Reconstruction of Document Structures中国科学技术大学 · 2023年



