DocLayNet
收藏资源简介:
DocLayNet是由IBM研究院瑞士Rueschlikon分部创建的大型人工标注文档布局分析数据集。该数据集包含80,863个来自多样数据源的手动标注页面,旨在代表布局的广泛变异性。每个PDF页面都提供了11个不同类别的标签边界框。DocLayNet还提供了部分页面进行双倍和三倍标注,以确定标注者间的一致性。数据集适用于训练深度学习模型,以提高文档转换中的布局检测和分割精度,特别是在处理复杂和多样化的布局时。
DocLayNet is a large-scale manually annotated document layout analysis dataset developed by the Rueschlikon division of IBM Research, Switzerland. This dataset comprises 80,863 manually annotated pages collected from diverse data sources, designed to represent the extensive variability of document layouts. Each PDF page is equipped with labeled bounding boxes across 11 distinct categories. Furthermore, DocLayNet provides double and triple annotations for a subset of pages to quantify inter-annotator agreement. This dataset is applicable for training deep learning models to enhance the accuracy of layout detection and segmentation in document conversion, particularly when handling complex and diverse document layouts.

- 1DocLayNet: A Large Human-Annotated Dataset for Document-Layout AnalysisIBM研究院瑞士Rueschlikon分部 · 2022年



