遇见数据集

Data set of the paper "Publishing an OCR ground truth data set for reuse in an unclear copyright setting"

收藏
Zenodo2021-05-12 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The data set consists of a METS file for each of the PDFs that were used for transcription and a directory data/page_xml that contains the transcriptions of the ground truth in PAGE-XML format. In parallel to the data set publication, a data paper will be published that contains a detailed description of the data set. As soon as it is published, we will link to it. The corresponding source code can be found here https://github.com/millawell/ocr-data/tree/1.1

提供机构:
Zenodo
创建时间:
2021-05-07
二维码
社区交流群
二维码
科研交流群
商业服务