dh-unibe/kurrent-hanse-xvi-test-lines-crop-noaugment
收藏资源简介:
该数据集名为kurrent-hanse-xvi-test-lines-crop-noaugment,是一个用于手写文本识别(HTR)的测试数据集,专注于历史文档中的Kurrent字体(一种德文手写体)文本行。数据集包含164个样本,全部位于训练分割中,总大小约为10.13 MB。数据通过pagexml-hf转换器从Transkribus PageXML格式转换而来,组织为parquet文件。每个样本包括图像、转录文本、行和区域的标识符、阅读顺序、坐标、基线信息以及文件名和项目名称等特征。数据集未进行数据增强,适用于图像到文本转换、转录任务,特别是基于Transformer的OCR(如TrOCR)模型的测试和评估。包含的项目涉及历史文档,如1505年的汉萨同盟相关记录。
The dataset named kurrent-hanse-xvi-test-lines-crop-noaugment is a test dataset for Handwritten Text Recognition (HTR), focusing on text lines in historical documents using Kurrent script (a German handwriting style). It contains 164 samples, all in the train split, with a total size of approximately 10.13 MB. The data was converted from Transkribus PageXML format using the pagexml-hf converter and is organized as parquet files. Each sample includes features such as image, transcribed text, line and region identifiers, reading order, coordinates, baseline information, filename, and project name. The dataset has no augmentation applied and is suitable for image-to-text conversion, transcription tasks, particularly for testing and evaluating Transformer-based OCR models like TrOCR. Included projects involve historical documents, such as records related to the Hanseatic League from 1505.




