OCR-D-GT-VD-SBB
收藏资源简介:
A ground truth (GT) dataset created within the OCR-D project and consisting of 348 pages extracted from historical documents pertaining to the "Verzeichnis der im deutschen Sprachraum erschienenen Drucke" (VD), all of which have been digitised by Staatsbibliothek zu Berlin – Berlin State Library (SBB). The data publication consists of 348 .xml files with transcriptions for 348 .tif facsimile image files. The image files pertain to 67 distinct works; four images were extracted from each of the 65 works; from two further works, 49 and 39 images respectively were extracted to create the GT. The dataset is complemented by a .csv file which contains a mapping between the identifiers used in this dataset and the unique identifiers used in the digitised collections of Staatsbibliothek zu Berlin – Berlin State Library, as well as a filelisting in .csv format. Data selection was performed within the OCR-D project at Staatsbibliothek zu Berlin – Berlin State Library. The project is funded by the German Research Foundation DFG, project grant no. 460675868. Ground truth data were established by a digitisation service provider and post-corrected by staff members of the Berlin State Library, data curation and publication was done by two members of the team of the research project "Mensch.Maschine.Kultur – Künstliche Intelligenz für das Digitale Kulturelle Erbe" at Staatsbibliothek zu Berlin – Berlin State Library. The research project was funded by the Federal Government Commissioner for Culture and the Media (BKM), project grant no. 2522DIG002.
本数据集为OCR-D项目打造的真实标注(ground truth, GT)数据集,包含从与《德语区出版印刷品目录》(Verzeichnis der im deutschen Sprachraum erschienenen Drucke,简称VD)相关的历史文献中提取的348页内容,所有文献均由柏林州立图书馆(Staatsbibliothek zu Berlin – Berlin State Library, SBB)完成数字化。该数据出版物包含348个XML格式转录文件,对应348张.tif格式数字化传真图像文件。这些图像文件隶属于67部不同作品:其中65部作品各提取4张图像,剩余2部作品分别提取49张与39张图像,以此构建该GT数据集。数据集还附带一份CSV文件,用于映射本数据集所用标识符与柏林州立图书馆数字化馆藏的唯一标识符之间的对应关系,同时包含一份CSV格式的文件列表。数据遴选工作由柏林州立图书馆的OCR-D项目完成。该项目由德国研究基金会(Deutsche Forschungsgemeinschaft, DFG)资助,项目编号为460675868。真实标注数据由数字化服务供应商生成,并经柏林州立图书馆工作人员后期校正;数据管理与发布工作由柏林州立图书馆下属研究项目"人·机器·文化——面向数字文化遗产的人工智能"(Mensch.Maschine.Kultur – Künstliche Intelligenz für das Digitale Kulturelle Erbe)团队的两名成员完成。该研究项目由联邦政府文化与媒体专员公署(Bundesbeauftragte für Kultur und Medien, BKM)资助,项目编号为2522DIG002。



