遇见数据集

Graphemic HTR ground truth dataset for an Early New High German transcription model (15th centruy)

收藏
Zenodo2026-02-06 更新2026-05-26 收录
官方服务:

资源简介:

This repository contains a set of training data for ATR models (Kraken). It contains 50 pages of ground truth as image files (jpg) and transcription files (PAGE xml). The ground truth contains 50 pages including 2,177 lines with 18,626 word tokens and 113,491 characters. Please refer to the README.md file for further information. A ground truth dataset following a graphemic transcription of the same data conntained within this repository may be found here: Diplomatic HTR-Ground Truth dataset . Authors The data in this repository was prepared and curated by Adam Juszczak (ORCiD: 0009-0000-5330-6183) and Frederik Skidzun (ORCiD: 0009-0002-7712-4207) of the Regesta Imperii - Regesta of Emperor Frederik III. License This dataset is made available under the CC-BY 4.0 license.

提供机构:
Zenodo
创建时间:
2026-02-06
二维码
社区交流群
二维码
科研交流群
商业服务