遇见数据集

The Caroline Minuscule Model (CMM) - Documentation, Dataset, Guidelines and Evaluation Results

收藏
Zenodo2025-12-19 更新2026-05-26 收录
官方服务:

资源简介:

This documentation describes the Caroline Minuscule Model, an automated handwriting recognition tool for transcribing manuscripts from the 8th to 11th centuries CE. It is a supplement to the chapter 5.2 "Case Study 2 Tim Geelhaar: A general ATR model for multiple purposes the Caroline Minuscule Model", in: Michael Schonhardt, Tobias Hodel, Jan Odstrcilik, Tim Geelhaar (eds.), Automated Text Recognition: Theory, Platforms, Best Practices (Bielefeld: BIUP, 2026), [not yet published]. This chapter outlines background, manuscript selection, guidelines, and evaluation for the model training. The supplement extends the chapter with a list all manuscripts used containing links to image sources. The ground truth data itself is provided in the dataset. Transcription guidelines establish basic rules for the model, while evaluation results are documented in an Excel file based on previously unseen manuscript pages. The model was trained on Transkribus and available there https://www.transkribus.org/model/latin-carolingian-minuscule. Please be advised that the naming convention has been updated to use "Caroline" instead of "Carolingian," as the script continued to be employed well beyond the Carolingian era.

提供机构:
Zenodo
创建时间:
2025-12-19
二维码
社区交流群
二维码
科研交流群
商业服务