dh-unibe/kurrent-hanse-xvi-test-rawxml-with-inference
收藏资源简介:
该数据集名为kurrent-hanse-xvi-test-rawxml-with-inference,是从jwidmer/kurrent-hanse-xvi-test数据集派生而来,并基于danameyer/kurrent-hanse-xvi-test-lines-with-inference数据集进行了增强,添加了原始XML推理内容。数据集包含6个样本,仅有一个训练集分割。数据特征包括图像(未解码模式)、XML内容(大字符串类型)、文件名(大字符串类型)、项目名(大字符串类型)和推理XML(大字符串类型,标注日期为2026年6月5日)。数据以parquet格式组织,按分割和项目名分片存储,总大小约为50.08 MB。数据集主要用于XML、PageXML、推理和手写文本识别(HTR)相关任务,许可证为MIT。使用示例展示了如何通过HuggingFace的datasets库加载整个数据集或特定分割。
This dataset is named kurrent-hanse-xvi-test-rawxml-with-inference, derived from the jwidmer/kurrent-hanse-xvi-test dataset and enriched with raw XML inference derived from danameyer/kurrent-hanse-xvi-test-lines-with-inference. It contains 6 samples across a single train split. Features include image (with decode false), xml_content (large_string), filename (large_string), project_name (large_string), and inference_xml_20260605_172616 (large_string). Data is organized in parquet shards by split and project name, with an approximate total size of 50.08 MB. The dataset is tagged for XML, PageXML, inference, and HTR (Handwritten Text Recognition) tasks, and is licensed under MIT. Usage examples demonstrate loading the entire dataset or specific splits via the HuggingFace datasets library.




