READ-BAD
收藏资源简介:
READ-BAD数据集由罗斯托克大学和维也纳工业大学联合创建,包含2036页来自不同时间和地点的档案文档图像。该数据集挑战了文本行分割方法,因其包含多样的页面布局和退化情况。创建过程中,从9个不同的欧洲档案馆收集了近2000份文档,并通过DigiTexx进行文本区域和基线的标注。READ-BAD数据集的应用领域主要集中在历史文档的布局分析,旨在解决文本行检测和分割的难题,特别是在处理复杂布局和多种退化情况时。
The READ-BAD dataset, jointly developed by the University of Rostock and Vienna University of Technology, comprises 2036 archival document images sourced from diverse time periods and geographical locations. This dataset presents substantial challenges to text line segmentation methods, as it features varied page layouts and multiple document degradation scenarios. During the dataset's construction, nearly 2000 documents were collected from nine distinct European archives, with text regions and baselines annotated via DigiTexx. The primary application scope of the READ-BAD dataset lies in layout analysis for historical documents, where it aims to address the core difficulties in text line detection and segmentation, especially when handling complex layouts and diverse degradation conditions.

- 1READ-BAD: A New Dataset and Evaluation Scheme for Baseline Detection in Archival Documents罗斯托克大学 · 2017年



