遇见数据集

Diplomatic HTR ground truth dataset for an Early New High German transcription model (15th century), Version 2

收藏
Zenodo2026-07-08 更新2026-08-01 收录
官方服务:

资源简介:

This repository contains a set of training data for ATR models (Kraken). It contains 50 pages of ground truth as image files (jpg) and transcription files (PAGE xml). The ground truth contains 50 pages including 2,177 lines with 18,626 word tokens and 110,618 characters. Please refer to the README.md file for further information. A ground truth dataset following a graphemic transcription of the same data conntained within this repository may be found here: Graphemic HTR-Ground Truth dataset . Authors The data in this repository was prepared and curated by Adam Juszczak (ORCiD: 0009-0000-5330-6183) of the BBAW / Regesta of Emperor Frederik III and Frederik Skidzun (ORCiD: 0009-0002-7712-4207) of the AdW Mainz / Regesta Imperii Online. License This dataset is made available under the CC-BY 4.0 license.

提供机构:
Zenodo
创建时间:
2026-07-08
二维码
社区交流群
二维码
科研交流群
商业服务