dh-unibe/kurrent-hanse-xvi-test-lines-augment
收藏资源简介:
该数据集名为kurrent-hanse-xvi-test-lines-augment,是通过Transkribus PageXML数据使用pagexml-hf转换器创建的。它包含492个样本,全部用于训练集,总大小约为33.52 MB。数据集主要用于图像到文本任务,特别是手写文本识别(HTR)、TrOCR模型训练和转录应用。每个样本包括图像字段(不进行解码)、文本内容、行和区域的元数据(如ID、阅读顺序、坐标、基线、增强类型)、文件名和项目名称。数据组织为Parquet分片格式,按分割和项目名称分类。标签涉及手写识别和转录,许可证为MIT。数据集包含一个历史文档项目:1505-02-10_Hanserezess,_Lübeck_Dienstag_nach_Scholastice_1505_(SAHST_Rep__2,_I_040-4),旨在支持历史文档的数字化和转录研究。
This dataset, named kurrent-hanse-xvi-test-lines-augment, was constructed from Transkribus PageXML data via the pagexml-hf converter. It contains 492 samples, all assigned to the training split, with a total size of approximately 33.52 MB. This dataset is primarily intended for image-to-text tasks, particularly handwritten text recognition (HTR), TrOCR model training and transcription applications. Each sample includes an undecoded image field, text content, metadata for lines and regions (including ID, reading order, coordinates, baseline, augmentation type), filename and project name. The data is organized in Parquet shard format, categorized by data split and project name. Labels are related to handwritten recognition and transcription, and the dataset is licensed under MIT. The dataset includes one historical document project: 1505-02-10_Hanserezess,_Lübeck_Dienstag_nach_Scholastice_1505_(SAHST_Rep__2,_I_040-4), which aims to support research on the digitization and transcription of historical documents.




