遇见数据集

Manually validated PageXML files for images in monography "Mémoire sur St Domingue par H ? M. Michel"

收藏
Zenodo2024-09-18 更新2026-05-26 收录
官方服务:

资源简介:

Transcription of monography "Mémoire sur St Domingue par H ? M. Michel", dating from 1797 and dealing on slavery in Haiti (103 pages in total). Transcription contains 61 pages in PageXML format, useful for training a handwritten text recognition model. The PageXML files were created by applying a Transkribus model (French Model 1, see https://readcoop.eu/model/french-general-model/, or the non-public The Text Titan I) on the images at https://europeana.transcribathon.eu/documents/story/?story=12733. The PageXML output was automatically corrected using the flat-text manual transcription available with these images, and the resulting PageXML files were manually validated. The software for automatically correcting OCR output using flat-text manual transcriptions (and hence adding a link between image and text not present in the flat-text files) has been developed in the AI4Culture project (https://pro.europeana.eu/project/ai4culture-an-ai-platform-for-the-cultural-heritage-data-space). Note: transcriptions for pages 21, 22, 34 and 58 are not present yet.

本数据集为1797年完成的专著"Mémoire sur St Domingue par H. ? M. Michel"的转录文本,该专著聚焦海地奴隶制问题,全书共计103页。其中61页的转录内容采用PageXML格式,可用于手写文本识别模型的训练。这些PageXML文件通过将Transkribus模型(法语模型1,详见https://readcoop.eu/model/french-general-model/,或非公开模型The Text Titan I)应用于https://europeana.transcribathon.eu/documents/story/?story=12733处的图像生成。随后借助配套图像提供的纯文本手动转录结果,对生成的PageXML输出内容进行自动校正,并对最终得到的PageXML文件开展人工校验。本数据集所使用的、基于纯文本手动转录结果自动校正光学字符识别(Optical Character Recognition, OCR)输出内容(同时可补充纯文本文件中缺失的图像与文本关联关系)的软件,由AI4Culture项目开发(详见https://pro.europeana.eu/project/ai4culture-an-ai-platform-for-the-cultural-heritage-data-space)。注:第21、22、34及58页的转录内容暂未提供。

提供机构:
Zenodo
创建时间:
2024-09-18
二维码
社区交流群
二维码
科研交流群
商业服务