Diplomatic HTR ground truth dataset for water damaged text lines (ENHG, 15th century)
收藏官方服务:
资源简介:
This repository contains a set of training data for ATR models (Kraken). It contains 50 pages of ground truth as image files (jpg) and transcription files (PAGE xml). The ground truth contains 50 pages including 2,258 lines with 19,087 word tokens and 112,281 characters. Please refer to the README.md file for further information. This dataset ist combinable with the Diplomatic HTR ground truth dataset for an Early New High German transcription model (15th century), Version 2. Authors The data in this repository was prepared and curated by Henriette Flohr (ORCiD: 0009-0005-6169-8320) and Jan Kunzek (ORCiD: 0009-0002-0430-7256) of the BBAW / Regesta of Emperor Frederik III. License This dataset is made available under the CC-BY 4.0 license.
提供机构:
Zenodo创建时间:
2026-07-19



