OCR-D-LAYOUT-GT-SBB
收藏资源简介:
A ground truth (GT) dataset created within the project “Robust and high-performance methods for layout analysis in OCR-D” and consisting of 276 pages extracted from historical documents pertaining to the "Verzeichnis der im deutschen Sprachraum erschienenen Drucke" (VD), all of which have been digitised by Staatsbibliothek zu Berlin – Berlin State Library (SBB). The data publication consists of 552 PAGE-XML files with transcriptions on line and region level and of 276 .tif facsimile image files from which the xml files resulted. The image files pertain to 35 distinct works; four images were extracted from 32 works each; three image files were extracted from one work; from two further works, 49 and 96 images respectively were extracted to create the GT. The dataset is complemented by a .csv file which contains a mapping between the identifiers used in this dataset and the unique identifiers used in the digitised collections of Staatsbibliothek zu Berlin – Berlin State Library, as well as a filelisting in .csv format. Data selection was performed within the project “Robust and high-performance methods for layout analysis in OCR-D” at Staatsbibliothek zu Berlin – Berlin State Library. The project is funded by the German Research Foundation DFG, project grant no. 517459941.



