OCR-D-LAYOUT-GT-DTA
收藏资源简介:
A ground truth (GT) dataset created within the DFG project “Robust and high-performance methods for layout analysis in OCR-D” that consists of 5,944 pages extracted from historical documents. 4,413 pages were taken from the Deutsches Textarchiv (DTA); for these pages, PAGE-XML files with regions and reading order are provided. Moreover, for the same images, PAGE-XML files with text lines are made available. In addition, further 963 images as well as PAGE-XML files containing graphic regions (graphicRegion) and 568 images as well as PAGE-XML files containing marginalia were processed and included in this data publication. The dataset is complemented by a filelisting in .csv format. Data selection was initially performed within the OCR-D project (https://ocr-d.de/en/). The project “Robust and high-performance methods for layout analysis in OCR-D” aims to improve the quality and robustness of layout analysis for historical documents within OCR-D and thus ensure their aptitude for mass digitisation. It received funding by the German Research Foundation (DFG), project grant no. 517459941.



