遇见数据集

Manually validated PageXML files for images in "Diario del soldato Bruno Celestino"

收藏
Zenodo2024-09-13 更新2026-05-26 收录
官方服务:

资源简介:

Transcribed diary of Italian soldier Bruno Celestino in World War I (52 pages in total) in PageXML format (pages 2 to 49 were transcribed). These files are useful for training a handwritten text recognition model. The PageXML files were created by applying Transkribus' Italian Handwriting M1 model (https://readcoop.eu/model/italian-general-model/) on the images at https://europeana.transcribathon.eu/documents/story/?story=110659, automatically correcting the output using the flat-text manual transcription available with these images, and manually validating the resulting PageXML files. The software for automatically correcting OCR output using flat-text manual transcriptions (and hence adding a link between image and text not present in the flat-text files) has been developed as part of the AI4Culture project (https://pro.europeana.eu/project/ai4culture-an-ai-platform-for-the-cultural-heritage-data-space).

本数据集为第一次世界大战意大利士兵布鲁诺·切莱斯蒂诺(Bruno Celestino)的转录日记,共计52页,采用PageXML(PageXML)格式,仅完成第2至49页的转录工作。该数据集可用于训练手写文本识别模型。 本数据集的PageXML文件生成流程如下:将Transkribus意大利手写体M1模型(https://readcoop.eu/model/italian-general-model/)应用于https://europeana.transcribathon.eu/documents/story/?story=110659 处的图像资源,再利用该图像配套的纯文本手动转录结果自动校正模型输出,并对最终生成的PageXML文件进行人工校验。 用于基于纯文本手动转录结果自动校正光学字符识别(Optical Character Recognition, OCR)输出(同时补充纯文本文件中缺失的图像与文本间关联)的软件,系AI4Culture项目(https://pro.europeana.eu/project/ai4culture-an-ai-platform-for-the-cultural-heritage-data-space)的开发成果之一。

提供机构:
Zenodo
创建时间:
2024-09-13
二维码
社区交流群
二维码
科研交流群
商业服务