dh-unibe/towerbooks-rawxml-test
收藏资源简介:
该数据集名为towerbooks-rawxml-test,是一个用于图像到文本转换、手写文本识别(HTR)、TrOCR模型训练和转录任务的测试数据集。它由Transkribus PageXML数据通过pagexml-hf转换器创建,包含4个样本,仅有一个训练分割。每个样本包括图像(以未解码格式存储)、xml_content(PageXML格式的文本内容)、文件名和项目名称。数据以parquet分片形式组织,按分割和项目名分类,便于在HuggingFace Hub上自动合并加载。数据集大小为约49.82 MB,许可证为MIT,适用于NLP和计算机视觉研究,特别是文档图像处理和转录任务。
This dataset, named towerbooks-rawxml-test, is a test dataset for image-to-text conversion, handwritten text recognition (HTR), TrOCR model training and transcription tasks. It was created from Transkribus PageXML data via the pagexml-hf converter, and contains 4 samples with only one training split. Each sample includes an image (stored in undecoded format), xml_content (text content in PageXML format), file name, and project name. The data is organized in parquet shards, categorized by split and project name, facilitating automatic merged loading on the Hugging Face Hub. The dataset has a size of approximately 49.82 MB, uses the MIT license, and is suitable for NLP and computer vision research, especially document image processing and transcription tasks.




