dh-unibe/kurrent-hanse-xvi-test-lines-crop-augment
收藏资源简介:
该数据集名为kurrent-hanse-xvi-test-lines-crop-augment,由Transkribus PageXML数据通过pagexml-hf转换器创建,专门用于手写文本识别(HTR)任务,特别是针对历史德文Kurrent手写体。数据集包含来自Hanse XVI测试集的492个样本,所有样本均属于训练集(train split)。每个样本包括裁剪和增强后的行图像、对应的文本转录、行和区域的元数据(如唯一标识符、阅读顺序、坐标、基线信息),以及文件名和项目名称。数据以parquet文件格式组织,便于高效加载和处理。该数据集适用于训练和评估图像到文本模型,如TrOCR等。
This dataset is named kurrent-hanse-xvi-test-lines-crop-augment. It was created from Transkribus PageXML data via the pagexml-hf converter, and is specifically designed for handwritten text recognition (HTR) tasks, particularly targeting historical German Kurrent script. The dataset contains 492 samples sourced from the Hanse XVI test set, with all samples belonging to the training split. Each sample includes cropped and augmented line-level images, corresponding text transcriptions, metadata for lines and regions (such as unique identifiers, reading order, coordinates, baseline information), as well as filenames and project names. The data is organized in Parquet file format to facilitate efficient loading and processing. This dataset is suitable for training and evaluating image-to-text models such as TrOCR.




