dh-unibe/towerbooks-line-test
收藏资源简介:
该数据集名为towerbooks-line-test,是通过pagexml-hf转换器从Transkribus PageXML数据创建的。它包含729个训练样本,用于图像到文本任务,特别是手写文本识别(HTR)和转录。数据特征包括图像、文本内容、行和区域的元数据(如ID、阅读顺序、坐标、基线、增强信息等),以及文件名和项目名称。数据以parquet文件格式组织,按项目和分割分片。适用于基于TroCR等模型的文本识别研究。
This dataset, named towerbooks-line-test, is created from Transkribus PageXML data via the pagexml-hf converter. It contains 729 training samples for image-to-text tasks, particularly Handwritten Text Recognition (HTR) and transcription. Its data features include images, text content, metadata for lines and regions (such as ID, reading order, coordinates, baselines, augmentation information, etc.), as well as filenames and project names. The data is organized in Parquet file format and sharded by project and data split. It is suitable for text recognition research based on models such as TroCR.




