遇见数据集

GT4HistOCR: Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin

收藏
Zenodo2021-02-21 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>GT4HistOCR</strong> contains ground truth for research in Optical Character Recognition (OCR) technology applied to historical printings in German Fraktur and Early Modern Latin. The ground truth comes in pairs of images of single printed lines as they appear in book pages (*.png) and their corresponding diplomatic transcriptions (*.gt.txt), which are UTF-8 strings preserving the character forms (glyphs) as much as possible within the UNICODE standard. These pairs of line images and their transcriptions can be directly used to train recognition models with, e.g., the open source OCR engines <em>OCRopy</em> or <em>Tesseract</em>. A total of 313,173 ground truth lines are provided. <strong>Please note that the subcorpora making up this collection used different transcription guidelines, so it is a bad idea to train a recognition model on the total collection! Rather train individual models for each subcorpus.</strong> Fur further information about the subcorpora, please see the README file and the accompanying publication. If these data are useful for you, please cite the accompanying publication: <pre>@article{springmann2018gt4hist, author = {Uwe Springmann and Christian Reul and Stefanie Dipper and Johannes Baiter}, title = {{Ground Truth for training {OCR} engines on historical documents in German Fraktur and Early Modern Latin}}, journal = {J. Lang. Technol. Comput. Linguistics}, volume = {33}, number = {1}, pages = {97--114}, year = {2018}, url = {https://jlcl.org/content/2-allissues/1-heft1-2018/jlcl_2018-1_5.pdf} }</pre>

<strong>GT4HistOCR</strong> 收录了面向德国Fraktur字体与早期近代拉丁语历史印刷品的光学字符识别(Optical Character Recognition,OCR)技术研究所需的真实标注数据。此类真实标注数据以成对形式提供:单印刷行的原始图像(取自书页,格式为*.png)及其对应的原样转录文本(格式为*.gt.txt);其中转录文本为UTF-8编码字符串,尽可能在Unicode标准范围内保留字符形态(字形)。此类行图像与转录文本对可直接用于训练识别模型,例如开源OCR引擎<em>OCRopy</em>与<em>Tesseract</em>。该数据集总计提供313,173条真实标注行数据。<strong>请注意:构成该数据集的各子语料库采用了不同的转录规范,因此直接在全数据集上训练识别模型并非恰当选择,建议针对每个子语料库分别训练专属模型。</strong>若需了解子语料库的更多信息,请参阅README文件及配套发表论文。若本数据集对您的研究有所帮助,请引用如下配套论文:<pre>@article{springmann2018gt4hist, author = {Uwe Springmann and Christian Reul and Stefanie Dipper and Johannes Baiter}, title = {{Ground Truth for training {OCR} engines on historical documents in German Fraktur and Early Modern Latin}}, journal = {J. Lang. Technol. Comput. Linguistics}, volume = {33}, number = {1}, pages = {97--114}, year = {2018}, url = {https://jlcl.org/content/2-allissues/1-heft1-2018/jlcl_2018-1_5.pdf} }</pre>

提供机构:
Zenodo
创建时间:
2018-08-12
二维码
社区交流群
二维码
科研交流群
商业服务