遇见数据集

OCR-D-LAYOUT-GT-SBB

收藏
Zenodo2026-04-29 更新2026-05-26 收录
官方服务:

资源简介:

A ground truth (GT) dataset created within the project “Robust and high-performance methods for layout analysis in OCR-D” and consisting of 276 pages extracted from historical documents pertaining to the "Verzeichnis der im deutschen Sprachraum erschienenen Drucke" (VD), all of which have been digitised by Staatsbibliothek zu Berlin – Berlin State Library (SBB). The data publication consists of 552 PAGE-XML files with transcriptions on line and region level and of 276 .tif facsimile image files from which the xml files resulted. The image files pertain to 35 distinct works; four images were extracted from 32 works each; three image files were extracted from one work; from two further works, 49 and 96 images respectively were extracted to create the GT. The dataset is complemented by a .csv file which contains a mapping between the identifiers used in this dataset and the unique identifiers used in the digitised collections of Staatsbibliothek zu Berlin – Berlin State Library, as well as a filelisting in .csv format. Data selection was performed within the project “Robust and high-performance methods for layout analysis in OCR-D” at Staatsbibliothek zu Berlin – Berlin State Library. The project is funded by the German Research Foundation DFG, project grant no. 517459941.

提供机构:
Zenodo
创建时间:
2026-04-29
二维码
社区交流群
二维码
科研交流群
商业服务