遇见数据集

HTR model for Latin and French Medieval Documentary Manuscripts (12th-15th)

收藏
Zenodo2023-01-18 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

This is the best HTR model for documentary Latin and French manuscripts presented in the paper: Sergio Torres Aguilar, Vincent Jolivet. <strong>Handwritten Text Recognition for Documentary Medieval<br> Manuscripts. </strong>2022. https://hal.science/hal-03892163 The model was trained on a charters and registers dataset from the Late-medieval period (12th-15th). The training and evaluation, entailing 1855 pages, 120k lines of text and almost 1M tokens, were conducted using three freely available ground-truth corpora : <strong>The Alcar-HOME database </strong>: https://zenodo.org/record/5600884 <strong>The e-NDP corpus </strong>: https://zenodo.org/record/7575693 <strong>The Himanis project </strong>: https://zenodo.org/record/5535306 This final model operates in a multilingual environment (Latin and Old French) and it is able to recognize several Latin script families (mostly <em>Textualis</em> and <em>Cursiva</em>) in documents produced in ca. 12th - 15th centuries. During the evaluation the models shows an accuracy of <strong>94.1%</strong> on the validation set and a CER (character error ratio) of about <strong>0.12</strong> to <strong>0.17</strong> on four external unseen datasets. A fine-tuning exercise using 10 ground-truth pages can raise these results to a CER between <strong>0.06</strong> to <strong>0.10</strong> respectively.

创建时间:
2023-01-18
二维码
社区交流群
二维码
科研交流群
商业服务