The Caroline Minuscule Model (CMM) - Documentation, Dataset, Guidelines and Evaluation Results
收藏资源简介:
This documentation describes the Caroline Minuscule Model, an automated handwriting recognition tool for transcribing manuscripts from the 8th to 11th centuries CE. It is a supplement to the chapter 5.2 "Case Study 2 Tim Geelhaar: A general ATR model for multiple purposes the Caroline Minuscule Model", in: Michael Schonhardt, Tobias Hodel, Jan Odstrcilik, Tim Geelhaar (eds.), Automated Text Recognition: Theory, Platforms, Best Practices (Bielefeld: BIUP, 2026), [not yet published]. This chapter outlines background, manuscript selection, guidelines, and evaluation for the model training. The supplement extends the chapter with a list all manuscripts used containing links to image sources. The ground truth data itself is provided in the dataset. Transcription guidelines establish basic rules for the model, while evaluation results are documented in an Excel file based on previously unseen manuscript pages. The model was trained on Transkribus and available there https://www.transkribus.org/model/latin-carolingian-minuscule. Please be advised that the naming convention has been updated to use "Caroline" instead of "Carolingian," as the script continued to be employed well beyond the Carolingian era.



