HTR Model Spanish Gothic Incunabula (HSMS)
收藏资源简介:
The Spanish Gothic Incunabula (HSMS) is conceived to be uploaded inside Transkribus platform (READ Coop) to perform a training and create an PyLaia model for the automated recognition of Spanish incunabula in Gothic script printed between 1472 and 1500. It can be used for post-incunabula (up to 1520).The transcription model follows the rules set by the Hispanic Seminary of Medieval Studies in 1977 (newest version). The rules applied are: Abbreviated words are expanded and the expanded text is enclosed between < >: q<ue> Superscripted letters are followed by a grave accent: q<u>i`en All ç and ñ are transcribed as ç and ñ No attemp has been made to normalize spacing All punctuations signs are kept All abbreviated nasals before b or p are transcribed as <n>. It is up to the editors if they should be changed into m. Abbreviated v' (tilde over v, or small slash v) that can be expanded as v<ir> or v<er> is expanded as v<er>. It is up to the editor if they should be changed to v<ir>. Tironian et is transcribed as & (ampersand) Pilcrows are transcribed as ¶ The model is built on 200 openings (verso-recto) drawn from 20 books printed by five different workshops form Sevile, Zaragoza, Burgos, Toledo and Pamplona. They Train Set consist 180152 words, distribuited over 24061 lines. The CER on the Train Set is 0.20 % and on the Validation Set 0.77 %.
西班牙哥特式古版印刷数据集(Spanish Gothic Incunabula, HSMS)旨在上传至Transkribus平台(READ Coop开发),用于模型训练并构建PyLaia模型,以自动识别1472年至1500年间以哥特体印刷的西班牙早期印刷书籍。该数据集亦可用于1520年及以前的后古版书时期文本识别。 本转录模型遵循西班牙中世纪研究学院(Hispanic Seminary of Medieval Studies)1977年发布的最新版规范,具体规则如下: 1. 缩写词需完成展开操作,展开后的文本需置于< >符号内,示例:q<ue> 2. 上标字母后需添加沉音符(grave accent),示例:q<u>i`en 3. 所有ç与ñ字符均原样转录 4. 未对文本间距进行标准化处理 5. 所有标点符号均予以保留 6. 所有位于b或p之前的缩合鼻音均转录为<n>,是否需将其替换为m由编辑自行决定 7. 可展开为v<ir>或v<er>的缩合符号v'(即v上方带有波浪符或带短斜杠的v),默认展开为v<er>,是否调整为v<ir>由编辑自行决定 8. 蒂罗尼安et缩写符(Tironian et)转录为&(and符号,ampersand) 9. 段落符号(pilcrows)转录为¶ 本模型的构建数据源自塞维利亚、萨拉戈萨、布尔戈斯、托莱多与潘普洛纳五家不同印刷工坊的20部印刷书籍,共计200个对开页(正反页)。训练集共包含180152个词,分布于24061行文本中。训练集上的字符错误率(Character Error Rate, CER)为0.20%,验证集上的字符错误率为0.77%。



